Information processing method and device based on large model, electronic equipment and medium

Through the information processing method based on big model, automatic crawling and analyzing web page contents is solved, and the problem of the existing technology is difficult to summarize forum-type web pages and multi-page reply information is achieved, and efficient and accurate content summary and user experience improvement are achieved.

CN120067472APending Publication Date: 2025-05-30BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510152135.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively summarize forum-type web page content containing a large amount of replies information, especially the inability to cover or summarize the replies information of subsequent multiple pages.

Method used

Using large-model-based information processing methods, through automated crawling and large-model analysis, target instruction information is obtained, including target paths and target topics, web page content is processed to generate relevant processing results, and page turn operations are supported to cover multi-page content.

Benefits of technology

It improves the speed and accuracy of generating processing results related to the topics that users pay attention to, and can cover or summarize content in subsequent multiple pages, improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067472A_ABST
    Figure CN120067472A_ABST
Patent Text Reader

Abstract

The invention provides an information processing method and device based on a large model, electronic equipment, a computer readable storage medium and a computer program product, and relates to the field of artificial intelligence, in particular to the technical field of large models, information processing and natural language processing. According to the implementation scheme, target instruction information is obtained, wherein the target instruction information comprises a target path and a target theme; obtaining first content in the target webpage based on the target path; processing the first content through a large model based on the target topic to obtain a first processing result; and in response to the received page turning instruction information and the obtained second content in the next page of the target webpage, processing the second content through the large model based on the target theme to obtain a second processing result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and in particular to the fields of large models, information processing, and natural language processing. Specifically, it relates to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for information processing based on a large model. Background Art

[0002] Artificial intelligence is a discipline that studies making computers simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and it has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing: Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0003] Human-computer interaction is a way for humans to interact with machines using natural language. With the continuous development of artificial intelligence technology, it has been realized that machines can understand the information output by humans, understand the inherent meaning in the information, and make corresponding feedback. In these operations, the accurate understanding of semantics, the speed of feedback, and the giving of corresponding opinions or suggestions all become factors affecting the smoothness of human-computer interaction.

[0004] The methods described in this section are not necessarily methods that have been previously envisioned or adopted. Unless otherwise specified, no method described in this section should be considered prior art solely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention

[0005] The present disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for information processing based on a large model.

[0006] According to one aspect of the present disclosure, there is provided a method for information processing based on a large model, including: obtaining target instruction information, where the target instruction information includes a target path and a target theme; obtaining first content in a target web page based on the target path; processing the first content through a large model based on the target theme to obtain a first processing result; in response to receiving paging instruction information and obtaining second content in the next page of the target web page, processing the second content through the large model based on the target theme to obtain a second processing result.

[0007] According to another aspect of the present disclosure, there is provided an information processing apparatus based on a large model, including: a first acquisition unit configured to acquire target instruction information, where the target instruction information includes a target path and a target theme; a second acquisition unit configured to acquire first content in a target web page based on the target path; a first processing unit configured to process the first content through the large model based on the target theme to obtain a first processing result; and a paging unit configured to, in response to receiving paging instruction information and acquiring second content in the next page of the target web page, process the second content through the large model based on the target theme to obtain a second processing result.

[0008] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the present disclosure.

[0009] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described in the present disclosure.

[0010] According to another aspect of the present disclosure, there is provided a computer program product including a computer program that implements the method described in the present disclosure when executed by a processor.

[0011] According to one or more embodiments of the present disclosure, through automated scraping and large model analysis, the speed and accuracy of generating processing results related to the user's concerned theme are improved; in addition, the system supports paging operations, can cover or summarize the content in subsequent multiple pages, and extract information matching the user's concerned theme therefrom, improving the user experience.

[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The drawings exemplarily illustrate embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary implementation manners of the embodiments. The illustrated embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0014] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein can be implemented according to an embodiment of the present disclosure;

[0015] Figure 2 Shows a flowchart of a large model-based information processing method according to an embodiment of the present disclosure;

[0016] Figure 3 Shows a flowchart of analyzing target web page content to obtain a processing result according to an embodiment of the present disclosure;

[0017] Figure 4 Shows a schematic diagram of instructions for generating a processing result including first summary information and second summary information according to an embodiment of the present disclosure;

[0018] Figure 5 Shows a block diagram of a large model-based information processing apparatus according to an embodiment of the present disclosure; and

[0019] Figure 6 Shows a block diagram of an exemplary electronic device capable of implementing the embodiments of the present disclosure. Detailed Description of the Embodiment

[0020] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0021] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.

[0022] In the description of various examples in the present disclosure, the terms used are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.

[0023] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0024] Figure 1FIG. shows a schematic diagram of an exemplary system 100 in which the various methods and apparatuses described herein may be implemented according to embodiments of the present disclosure. Referring to Figure 1 , the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 may be configured to execute one or more applications.

[0025] In embodiments of the present disclosure, the server 120 may run one or more services or software applications that enable methods for performing information processing based on large models.

[0026] In certain embodiments, the server 120 may also provide other services or software applications, which may include non-virtual environments and virtual environments. In certain embodiments, these services may be provided as web-based services or cloud services, for example, provided to users of the client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) model.

[0027] In Figure 1 the configuration shown, the server 120 may include one or more components that implement the functions performed by the server 120. These components may include software components, hardware components, or combinations thereof that may be executed by one or more processors. Users operating the client devices 101, 102, 103, 104, 105, and / or 106 may, in turn, utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that various different system configurations are possible, which may differ from the system 100. Therefore, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.

[0028] Users may use the client devices 101, 102, 103, 104, 105, and / or 106 to summarize web page information. The client device may provide an interface that enables the user of the client device to interact with the client device. The client device may also output information to the user via the interface. Although Figure 1 only six client devices are depicted, those skilled in the art will be able to understand that the present disclosure may support any number of client devices.

[0029] Client devices 101, 102, 103, 104, 105, and / or 106 can include various types of computing devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computing devices can run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux, or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT WindowsMobile OS, iOS, Windows Phone, Android. Portable handheld devices can include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices can include head-mounted displays (such as smart glasses) and other devices. Gaming systems can include various handheld gaming devices, Internet-enabled gaming devices, etc. Client devices are capable of executing various different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.

[0030] Network 110 can be any type of network known to those skilled in the art, which can support data communication using any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, token ring, wide area network (WAN), the Internet, virtual network, virtual private network (VPN), intranet, extranet, blockchain network, public switched telephone network (PSTN), infrared network, wireless network (such as Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0031] Server 120 can include one or more general-purpose computers, dedicated server computers (such as PC (personal computer) servers, UNIX servers, midrange servers), blade servers, mainframes, server clusters, or any other suitable arrangement and / or combination. Server 120 can include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for the server). In various embodiments, server 120 can run one or more services or software applications that provide the functions described below.

[0032] The computing unit in server 120 can run one or more operating systems including any of the above operating systems and any commercially available server operating systems. Server 120 can also run any one of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0033] In some embodiments, server 120 can include one or more applications to analyze and combine data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 can also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.

[0034] In some embodiments, server 120 can be a server of a distributed system, or a server combined with a blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability existing in traditional physical hosts and virtual private server (VPS) services.

[0035] System 100 can also include one or more databases 130. In certain embodiments, these databases can be used to store data and other information. For example, one or more of databases 130 can be used to store information such as target instruction information, web page content, etc. Databases 130 can reside in various locations. For example, the databases used by server 120 can be local to server 120, or can be remote from server 120 and can communicate with server 120 via a network-based or dedicated connection. Databases 130 can be of different types. In certain embodiments, the databases used by server 120 can be relational databases, for example. One or more of these databases can store, update, and retrieve data to and from the databases in response to commands.

[0036] In certain embodiments, one or more of databases 130 can also be used by applications to store application data. The databases used by applications can be different types of databases, such as key-value repositories, object repositories, or conventional repositories supported by a file system.

[0037] Figure 1The system 100 can be configured and operated in various ways to enable the application of various methods and devices described in this disclosure.

[0038] Currently, there are many tools that can summarize web page information. Generally, tools based on traditional deep learning methods can only summarize the main part of a web page (such as the title and the information in the front part of the home page), and are ineffective for forum websites. Forum websites usually start with a topic, and the content consists of the topic, header information, and responder information, and the responder information often occupies most of the forum content. Currently, web page information summarization tools cannot cover or summarize the reply information on subsequent multiple pages, even if these replies contain a lot of valuable content.

[0039] Therefore, the method according to this disclosure provides an information processing method based on a large model. Figure 2 A flowchart of an information processing method based on a large model according to an embodiment of this disclosure is shown, as Figure 2 shown, the method 200 includes: obtaining target instruction information, where the target instruction information includes a target path and a target topic (step 210); obtaining first content in a target web page based on the target path (step 220); processing the first content through a large model based on the target topic to obtain a first processing result (step 230); in response to receiving paging instruction information and obtaining second content in the next page of the target web page, processing the second content through the large model based on the target topic to obtain a second processing result (step 240).

[0040] According to an embodiment of this disclosure, through automated crawling and large model analysis, the speed and accuracy of generating processing results related to the user's concerned topic are improved; in addition, the system supports paging operations, can cover or summarize the content in subsequent multiple pages, and extract information that matches the user's concerned topic, improving the user experience.

[0041] In some embodiments, the target path may include web page address information, and the web page address information may include the address url of the web page, etc. The web page can be located according to the web page url to extract the desired data. After obtaining the target path, all relevant content on the corresponding web page can be accessed and crawled according to the target path provided by the user. The crawling method includes but is not limited to using crawler tools, and the crawled content includes but is not limited to information such as text, comments, etc. extracted from the web page, such as titles, texts, replies, and comments, etc.

[0042] In some examples, the target path can be the link address of a specific forum post provided by the user, and this link points to the specific web page content that the user hopes to analyze.

[0043] In some examples, a web page information acquisition tool (such as Crawl4AI) can automatically identify and parse web page elements based on the capabilities of large models, simplifying the web crawling and data extraction processes. Such tools are based on an asynchronous architecture, can efficiently process multiple pages to quickly scrape the required data, and support multiple output formats (including JSON, HTM, Markdown) to meet the data requirements of different scenarios. For example, through an API interface, the path of the web page (target web page or next page) and the language-based scraping rules are passed to the relevant large model-based web page information scraping tool, allowing the large model to scrape the web page information according to the rule instructions and return the organized web page information content to the system through the API.

[0044] In some examples, if the large model returns a web page information value of 200, it indicates that the web page exists and is valid, and subsequent operations can be performed; if the web page information value is not 200, it indicates that the page does not exist, such as a path error or the current page is the last page of the post, and a prompt message is returned and the conversation ends.

[0045] In some embodiments, the target topic can include specific topics or keywords of interest to the user, and the system can screen and analyze the web page content based on the target topic.

[0046] In the above embodiments, if the system receives a page turning instruction message, it can continue to access the content of the next page using the same web page content scraping method based on the previously saved memory variables (such as the URL and Topic mentioned above) to ensure context consistency. After obtaining the web page content of the next page, the large model can process this content using the same processing method as described above to output a second processing result.

[0047] According to some embodiments, the obtaining of the second processing result by processing the second content in the next page of the target web page based on the target topic through the large model in response to receiving the page turning instruction message and obtaining the second content includes: in response to receiving the page turning instruction message, obtaining the path information corresponding to the next page of the target web page based on the target path; and obtaining the second content based on the path information.

[0048] In some examples, the path information corresponding to the next page can be generated according to the target path in the previously saved memory variables (such as the URL and Topic mentioned above). For example, the URL corresponding to the target page is XXXXX / page=1, and then when the "page turning instruction" (i.e., the next page) is obtained, page+1 is used to obtain the new URL = XXXXX / page=2. Then, all relevant content on the corresponding web page can be accessed and scraped based on the new URL.

[0049] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information and other processing comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0050] According to some embodiments, obtaining the target instruction information includes: in response to determining that the first input information including the target path and the target topic is received, obtaining the target instruction information based on the first input information.

[0051] In the above embodiment, the first input information may include, but is not limited to, information such as text, voice, and pictures. The first input information is used for the system to feedback corresponding reply information based on the input information. For example, when the first input information is "What's the weather like today?", the corresponding reply content may include the weather conditions of the current date and region. In some examples, when input in voice form, the voice can be converted into text through voice recognition technology to obtain the first retrieval information.

[0052] In the above embodiment, after receiving the first input information, the system will check the content of the information to confirm whether it contains two key factors, namely the target path (URL) and the target topic (Topic). This step can ensure that the system can clearly identify the specific web page that the user wants to process and the topic of concern. Once it is determined that the first input information contains the target path and the target topic, it means that the user expects to execute the information processing method described in the embodiments of the present disclosure (i.e., intention recognition), such as scraping web page content, target web page and other content processing, rather than a simple reply operation such as the weather information described above.

[0053] According to some embodiments, obtaining the target instruction information includes: in response to determining that the first input information that does not include at least one of the target path and the target topic is received, outputting a prompt message; and in response to determining that the second input information including at least one of them is received, obtaining the target instruction information based on the first input information and the second input information.

[0054] In the above embodiment, when the system detects that the first input information is incomplete, it can output a prompt message for inquiry to guide the user to provide the missing information. This kind of prompt message is usually presented in the form of a dialog box, command line prompt or other interaction methods, asking the user to supplement the missing target path or target topic. For example, the system may prompt the user "You have not provided the target path, please enter the URL of the post" or "Please specify the topic you are interested in", etc. After the user provides supplementary information according to the prompt, these new inputs are called "second input information". The second input information may include at least one of the previously missing target path or target topic information. At this time, the system can combine the first input information and the second input information to form complete target instruction information.

[0055] Therefore, this supplementary mechanism can ensure that even if the information initially provided by the user is incomplete, the system can determine the specific operation the user expects to perform and obtain all the required information through follow-up questions, thereby ensuring the smooth progress of the entire processing flow. This can not only improve the fault tolerance of the system but also, based on complete instruction information, enhance the efficiency and accuracy of the system in analyzing and processing web page content.

[0056] According to some embodiments, the target topic includes: a first topic and a second topic, and each second topic is a subcategory of the corresponding first topic in the first topic. Based on the target topic, processing the first content through a large model to obtain a first processing result includes: respectively based on the first topic and the second topic, summarizing the content of the first content through the large model to obtain corresponding third and fourth processing results; the first processing result includes: the third processing result and the fourth processing result.

[0057] In some embodiments, the content in the target web page includes body content and comment content. The target topic includes at least one first topic and at least one second topic, where each second topic in the at least one second topic is a subcategory of the corresponding first topic in the at least one first topic. That is, the second topic is a fine-grained classification of the corresponding first topic.

[0058] In some embodiments, the content of the target web page includes body content and comment content. The body content is the main content part of the post, usually including a title, body text, and possibly pictures or other multimedia elements; the comment content is the replies and discussions made by users under the post, usually containing a large amount of interactive information and different viewpoints. On the other hand, the target topic includes at least one first topic and at least one second topic. The first topic can refer to a coarse-grained topic or keyword (i.e., the parent category of the second topic) for summarizing the main content of the web page. For example, "XX product", etc.; the second topic can refer to the fine-grained classification of the first topic, providing more specific concerns. For example, under the first topic of "XX product", the second topic can be "battery life", "processor speed", etc.

[0059] Therefore, in some embodiments, respectively based on the first topic and the second topic, summarizing the content of the first content through the large model to obtain corresponding third and fourth processing results includes: based on the first topic, summarizing the body content through the large model to obtain corresponding third processing results; based on the second topic, summarizing the comment content through the large model to obtain the fourth processing result.

[0060] It is understandable that the first topic and the second topic here can be the same or different. That is, the fine-grained classification of the first topic can include itself, and there is no limitation here.

[0061] According to some embodiments, as Figure 3 shown, processing the first content based on the target topic through a large model to obtain a first processing result (method 300) includes: performing content summarization on the body content through a large model based on the at least one first topic to obtain a third processing result (step 310); performing content summarization on the comment content through a large model based on the at least one second topic to obtain a fourth processing result (step 320); and using the third processing result and the fourth processing result together as the first processing result (step 330).

[0062] In the above example, the system can first analyze and summarize the body content of the web page according to the provided first topic using a large model. The goal of this step is to extract the core information related to the first topic from a macro level. For example, if the first topic is "XX product", the system will extract the information related to this product from the body text and generate a summary as the third processing result. Then, the system can deeply analyze and summarize the comment content according to the second topic. Since the second topic is a fine-grained classification of the first topic, this process can extract more specific and detailed discussion points. For example, for the second topic of "battery life", the system will screen the specific discussions and user feedback on the battery life of the product in the comments and summarize them to generate the fourth processing result. Finally, the system can integrate the third processing result (summary of the body content) and the fourth processing result (summary of the comment content) to form a comprehensive first processing result. That is to say, the first processing result can not only include the core information in the web page body, but also cover the specific discussions and user feedback related to each fine-grained topic in the comments. In this way, the system can provide users with a comprehensive and structured summary to help them quickly understand the main content of the web page and its related detailed discussions.

[0063] Through the above embodiments, for reply information including a large amount of excellent content such as in a forum, information processing such as screening, aggregation, and summarization can be performed more accurately, realizing the automation of finding the most concerned content in a large number of reply posts.

[0064] According to some embodiments, the first content in the target web page includes body content. The content summarization of the first content based on the first theme and the second theme respectively through a large model to obtain corresponding third and fourth processing results includes: in response to determining that the body content is associated with the first theme, content summarization of the first content is performed based on the first theme and the second theme respectively through the large model to obtain the third processing result and the fourth processing result.

[0065] In some embodiments, processing the first content based on the target theme through a large model to obtain a first processing result includes: before obtaining the third processing result and the fourth processing result, determining whether the body content is associated with one or more of the at least one first theme; in response to determining that the body content is associated with one or more of the first themes, obtaining the third processing result and the fourth processing result.

[0066] In some examples, after obtaining the web page content, it is possible to first determine whether the main content (i.e., the body content) in the target web page is associated with the topic that the user is interested in, and then, when determining that the relevance strength is relatively high, process the web page content, etc., to output a first processing result. For example, a large model can be used to determine whether the target web page is associated with the topic that the user is interested in.

[0067] In some embodiments, it can be understood that only when the body content is associated with at least one first theme, the summarization of the body content and the comment content can accurately reflect the topic that the user is interested in and provide more valuable information. Therefore, before obtaining the third processing result and the fourth processing result, it is necessary to analyze the body content through a large model to determine whether it is associated with the first theme specified by the user. This process is a prerequisite for ensuring the pertinence and effectiveness of the subsequent content summarization. For example, the large model can perform semantic analysis on the body content, identify the topics therein, compare them with the first themes provided by the user, and then evaluate the degree of association between the body content and each first theme based on semantic similarity or other algorithms. If the body content has a strong association with at least one first theme, it can be considered that the content meets the user's focus, and then the body content and the comment content can be summarized according to the above method to obtain the third processing result and the fourth processing result.

[0068] According to some embodiments, the first content in the target web page includes body content. The target topics include: a first topic and a second topic, where each of the second topics is a subcategory of the corresponding first topic in the first topic. The processing of the first content by the large model based on the target topic to obtain a first processing result includes: in response to determining that the body content is not associated with the first topic, summarizing the body content by the large model to obtain a fifth processing result. The first processing result includes the fifth processing result.

[0069] In some embodiments, the processing of the first content by the large model based on the target topic to obtain a first processing result includes: before obtaining the third processing result and the fourth processing result, determining whether the body content is associated with one or more of the at least one first topic; in response to determining that the body content is not associated with any of the at least one first topic, summarizing the body content by the large model to obtain a fifth processing result; and using the fifth processing result as the first processing result.

[0070] In some embodiments, as described above, before obtaining the third processing result and the fourth processing result, it is necessary to analyze the body content through the large model to determine whether it is associated with the first topic specified by the user. This process is a prerequisite for ensuring the pertinence and effectiveness of subsequent content summarization. If the body content is not associated with any of the at least one first topic, the body content can be summarized through the large model. This summary is not based on a specific first topic, but rather generalizes and refines the main content, core viewpoints, and important information of the entire body content, for example, to generate a fifth processing result. In this way, computing resources can be saved, that is, when the body content is not associated with the first topic, there is no need to further process its comment content. Moreover, when the body content fails to match the first topic specified by the user, the system can still provide a comprehensive summary of the body content to ensure that the user can understand the basic information of the web page, thereby improving the practicality of the system and enhancing the user experience.

[0071] According to some embodiments, in response to receiving paging instruction information and obtaining second content in the next page of the target web page, the processing of the second content by the large model based on the target topic to obtain a second processing result includes: in the case where it is determined that the body content is associated with the first topic, in response to receiving paging instruction information and obtaining the second content, processing the second content by the large model based on the target topic to obtain the second processing result.

[0072] In some embodiments, when it is determined that the body content is associated with one or more first topics, the user can issue a page-turning instruction through a certain interaction method, such as clicking the "Next Page" button or sending a specific command (such as sending "What about the next page"). At this time, the system will retain the previously saved memory variables (such as the target path, the target topic, and the first processing result) when receiving the page-turning request, because these variables can help the system remember the current context, so that it can continue to accurately process the relevant content after page-turning, ensuring the continuity and consistency of information processing, and ensuring that the content on the next page is also associated with the current topic.

[0073] In some examples, the system can access the next page of the target web page according to the previously saved target path and other relevant information. If the next page exists, continue to obtain the content on the next page of the target web page; otherwise, the system can prompt the user that the page does not exist or cannot be accessed. Here, the method of obtaining the content on the next page of the target web page is the same as the method of obtaining the content on the target web page, which will not be elaborated here.

[0074] According to some embodiments, the second content includes comment content. The processing of the second content by the large model based on the target topic to obtain a second processing result includes: summarizing the comment content by the large model based on the second topic to obtain the second processing result.

[0075] In some embodiments, when the user issues a page-turning instruction and successfully obtains the content of the next page, if the content of the next page contains comments, the comment content can be analyzed and summarized based on the second topic provided by the user. Here, the step of summarizing the comment content by the large model based on the second topic to obtain the second processing result is the same as the step of obtaining the fourth processing result for the target web page, which will not be elaborated here.

[0076] In this way, whether it is the content of a single page or the comments on multiple pages, comprehensive and targeted web page information can be provided for the user, so that during the process of the user browsing the content of multiple pages of the forum, the quality, matching degree, and accuracy of the content processing results closely related to the topic concerned by the user can be improved.

[0077] According to some embodiments, the comment content includes multiple first comment messages, wherein the fourth processing result includes a first summary message and a second summary message. The first summary message is obtained by summarizing the content of the first comment messages; and the second summary message is obtained by summarizing the content of second comment messages, wherein the second comment messages are the first comment messages associated with the second topic.

[0078] In some embodiments, the first summary information is the summary information obtained by summarizing the content of the plurality of first comment information; the second summary information is the summary information obtained by summarizing the second comment information, where the second comment information is the first comment information among the plurality of first comment information that is associated with the corresponding second topic in the at least one second topic.

[0079] In some embodiments, the comment section of a target web page typically contains multiple comment information, and these comments may cover different topics, viewpoints, and feedback. Specifically, the first summary information can be an overall summary of all comments, which can cover the main viewpoints and discussion points of all comments to provide a comprehensive perspective. For example, assuming that users are concerned about the discussion of a certain product, the first summary information may include that "most users are satisfied with the overall performance of the product, but some users also mention that the battery life is short". The second summary information can be a summary of the comments associated with a specific second topic (i.e., the first comment information, which can include one or more comment information), which focuses on the fine-grained topics specified by the user and provides more specific information, that is, it can provide a detailed discussion summary for a specific topic, enabling users to deeply understand the specific aspects they are most concerned about. For example, if one of the second topics is "battery life", the second summary information may include that "in the discussion about battery life, many users mentioned that the actual battery life during use is about 8 hours, and it is recommended to improve the battery management function".

[0080] Therefore, by summarizing all comment information at different levels to obtain the first summary information and the second summary information, it can ensure that users can not only obtain comprehensive information but also deeply understand the discussions and feedback on specific aspects.

[0081] According to some embodiments, the second summary information further includes at least one of the following items: the serial number of the second comment information in the first comment information, and the association relationship between the second comment information and the corresponding second topic. That is, the second summary information can include: the serial number of the second comment information in the plurality of first comment information, and / or the association relationship between the second comment information and the corresponding second topic.

[0082] In some embodiments, the comment section of a target web page typically contains multiple comment information. When there are multiple comment information, when summarizing the content of the corresponding comment information to generate the corresponding summary information, the original comment serial number (level) corresponding to the summary information can be marked, such as the 3rd comment, the 5th - 6th comments. In this way, it is convenient for users to trace back the corresponding information and improves the user experience.

[0083] In some embodiments, when summarizing the content of the corresponding comment information to generate the corresponding summary information, it is also possible to identify which second topic the summary information is associated with that the user is concerned about. In this way, it is convenient for the user to clarify the key points of the information and trace back the corresponding information, improving the user experience.

[0084] Figure 4 FIG. shows a schematic diagram of instructions for generating a processing result including first summary information and second summary information according to an embodiment of the present disclosure. As Figure 4 shown, the tasks are clearly defined in the instructions: 1) It is necessary to comprehensively summarize the comment content in the comment table (refer to 401); 2) Screen out the comments associated with the concerned topic from the comment table (refer to 402). And when screening out the comments associated with the concerned topic from the comment table, the format and content of the output summary information are further clarified in the instructions, including the reply serial number (refer to 403), content summary, and the matching concern points (i.e., the associated second topic, refer to 404).

[0085] In some embodiments, when implementing corresponding multiple operations through a large model, the large model for implementing the multiple operations can be the same large model or different large models, which is not limited herein.

[0086] In some embodiments, the large model can be a knowledge-enhanced large language model for dialogue (such as ERNIE bot, etc.), and the large model is trained with a vast amount of knowledge resources and dialogue data (such as including web data exceeding trillions, billions of search data, hundreds of millions of image data, billions of voice requests per day, text requests exceeding 50 billion times, and factual knowledge exceeding 550 billion).

[0087] Therefore, applying this type of model as the large model can not only directly process chat-type dialogue information, but also directly generate reply information for logical reasoning-type, common sense-type, and image generation-type dialogue information, capable of further improving the generation efficiency while generating higher-quality reply information.

[0088] According to an embodiment of the present disclosure, there is also provided an information processing device based on a large model. Figure 5 FIG. shows a schematic diagram of an information processing device 500 based on a large model according to an embodiment of the present disclosure. As Figure 5As shown, the device 500 includes: a first acquisition unit 510 configured to acquire target instruction information, where the target instruction information includes a target path and a target topic; a second acquisition unit 520 configured to acquire first content in a target web page based on the target path; a first processing unit 530 configured to process the first content through a large model based on the target topic to obtain a first processing result; and a paging unit 540 configured to, in response to receiving paging instruction information and acquiring second content in the next page of the target web page, process the second content through the large model based on the target topic to obtain a second processing result.

[0089] According to some embodiments, the first acquisition unit 510 includes: a unit configured to, in response to determining that first input information including the target path and the target topic is received, acquire the target instruction information based on the first input information.

[0090] According to some embodiments, the first acquisition unit 510 further includes: a unit configured to, in response to determining that the first input information that does not include at least one of the target path and the target topic is received, output prompt information; and a unit configured to, in response to determining that second input information including at least one of them is received, acquire the target instruction information based on the first input information and the second input information.

[0091] According to some embodiments, the target topic includes: a first topic and a second topic, where each second topic is a subcategory of the corresponding first topic in the first topic. The first processing unit 530 includes: a unit configured to respectively perform content summarization on the first content through a large model based on the first topic and the second topic to obtain corresponding third processing results and fourth processing results. The first processing result includes: the third processing result and the fourth processing result.

[0092] According to some embodiments, the first content in the target web page includes body content. The unit configured to respectively perform content summarization on the first content through a large model based on the first topic and the second topic to obtain corresponding third processing results and fourth processing results includes: a unit configured to, in response to determining that the body content is associated with the first topic, respectively perform content summarization on the first content through the large model based on the first topic and the second topic to obtain the third processing result and the fourth processing result.

[0093] According to some embodiments, the first content in the target web page includes the body content, and the target theme includes: a first theme and a second theme, and each of the second themes is a subcategory of the corresponding first theme in the first theme. The first processing unit 530 includes: a unit configured to, in response to determining that the body content is not associated with the first theme, summarize the body content through the large model to obtain a fifth processing result. The first processing result includes the fifth processing result.

[0094] According to some embodiments, the paging unit 540 includes: a unit configured to, when it is determined that the body content is associated with the first theme, in response to receiving paging instruction information and obtaining the second content, process the second content through the large model based on the target theme to obtain the second processing result.

[0095] According to some embodiments, the second content includes comment content. The paging unit 540 includes: a unit configured to summarize the comment content through the large model based on the second theme to obtain the second processing result.

[0096] According to some embodiments, the comment content includes multiple first comment messages, and the fourth processing result includes first summary information and second summary information. The first summary information is obtained by summarizing the first comment messages; and the second summary information is obtained by summarizing second comment messages, where the second comment messages are the first comment messages associated with the second theme.

[0097] According to some embodiments, the second summary information further includes at least one of the following items: the serial number of the second comment message in the first comment messages, and the association relationship between the second comment message and the corresponding second theme.

[0098] According to some embodiments, the paging unit 540 includes: in response to receiving the paging instruction information, obtaining the path information corresponding to the next page of the target web page based on the target path; and obtaining the second content based on the path information.

[0099] Here, the operations of the above units 510-540 of the information processing device 500 based on the large model are respectively similar to the operations of steps 210-240 described above, and will not be elaborated here.

[0100] According to an embodiment of the present disclosure, there is also provided an electronic device, a readable storage medium, and a computer program product.

[0101] Reference Figure 6, a block diagram of an electronic device 600 that can be a server or a client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0102] As Figure 6 shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0103] Multiple components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, an output unit 607, a storage unit 608, and a communication unit 609. The input unit 606 can be any type of device that can input information into the electronic device 600. The input unit 606 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device, and can include, but are not limited to, a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 607 can be any type of device that can present information, and can include, but are not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 608 can include, but are not limited to, magnetic disks, optical disks. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include, but are not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0104] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as method 200. For example, in some embodiments, method 200 may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of method 200 described above may be executed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute method 200 in any other suitable manner (e.g., by means of firmware).

[0105] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0106] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0107] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0108] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0109] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0110] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0111] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0112] Although embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, the steps can be executed in an order different from that described in this disclosure. Further, various elements in the embodiments or examples can be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein can be replaced by equivalent elements that emerge after this disclosure.

Claims

1. An information processing method based on a large model, comprising: Acquire target instruction information, wherein the target instruction information includes a target path and a target subject; Acquire first content in a target webpage based on the target path; Based on the target topic, the first content is processed by a large model to obtain a first processing result; In response to receiving the page turning instruction information and acquiring the second content in the next page of the target webpage, the second content is processed by the large model based on the target theme to obtain a second processing result.

2. The method of claim 1, wherein: The acquiring target instruction information comprises: In response to determining that first input information including the target path and the target subject is received, the target instruction information is acquired based on the first input information.

3. The method of claim 2, wherein: The acquiring target instruction information comprises: In response to determining that the first input information that does not include at least one of the target path and the target subject is received, outputting prompt information; and In response to determining that the second input information including the at least one item is received, the target instruction information is acquired based on the first input information and the second input information.

4. The method of claim 1, wherein: The target topics include: a first topic, and a second topic, wherein each of the second topics is a subcategory of a corresponding first topic in the first topic, and wherein, The processing of the first content by using a large model based on the target theme to obtain a first processing result includes: Based on the first theme and the second theme respectively, the first content is summarized by a large model to obtain a third processing result and a fourth processing result corresponding to each other; The first processing result includes: the third processing result and the fourth processing result.

5. The method of claim 4, wherein: The first content in the target webpage includes text content, wherein the first content is summarized by a large model based on the first theme and the second theme respectively to obtain a one-to-one corresponding third processing result and a fourth processing result, including: In response to determining that the main text content is associated with the first topic, the first content is summarized through the macro model based on the first topic and the second topic respectively to obtain the third processing result and the fourth processing result.

6. The method according to claim 1 or 5, wherein: The first content in the target webpage includes text content, wherein the target topic includes: a first topic, and a second topic, wherein each of the second topics is a subcategory of a corresponding first topic in the first topics, and wherein, The processing of the first content by using a large model based on the target theme to obtain a first processing result includes: In response to determining that the main text content is not associated with the first topic, summarizing the main text content by using the macro model to obtain a fifth processing result; Among them, the first processing result includes the fifth processing result.

7. The method according to claim 4 or 5, wherein: In response to receiving the page turning instruction information and acquiring the second content in the next page of the target webpage, based on the target theme, processing the second content by the large model to obtain a second processing result includes: When it is determined that the main text content is associated with the first theme, in response to receiving page turning instruction information and acquiring the second content, the second content is processed by the large model based on the target theme to obtain the second processing result.

8. The method according to any one of claims 4 to 7, wherein: The second content includes comment content, and wherein, based on the target topic, processing the second content by the macro model to obtain a second processing result includes: Based on the second topic, the review content is summarized by the large model to obtain the second processing result.

9. The method of claim 8, wherein: The comment content includes a plurality of first comment information, wherein the fourth processing result includes first summary information and second summary information, wherein, The first summary information is obtained by summarizing the content of the first comment information; and The second summary information is obtained by summarizing the content of second comment information, wherein the second comment information is the first comment information associated with the second topic.

10. The method of claim 9, wherein: The second summary information further includes at least one of the following items: a sequence number of the second comment information in the first comment information, and an association relationship between the second comment information and the corresponding second topic.

11. The method of claim 1, wherein: In response to receiving the page turning instruction information and acquiring the second content in the next page of the target webpage, based on the target theme, processing the second content by the large model to obtain a second processing result includes: In response to receiving the page turning instruction information, acquiring path information corresponding to the next page of the target web page based on the target path; and Based on the path information, the second content is acquired.

12. An information processing device based on a large model, comprising: A first acquisition unit is configured to acquire target instruction information, wherein the target instruction information includes a target path and a target subject; A second acquisition unit, configured to acquire first content in a target webpage based on the target path; A first processing unit is configured to process the first content through a large model based on the target theme to obtain a first processing result; The page turning unit is configured to, in response to receiving the page turning instruction information and acquiring the second content in the next page of the target web page, process the second content through the large model based on the target theme to obtain a second processing result.

13. An electronic device comprising: at least one processor; as well as a memory communicatively coupled to the at least one processor; in The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 11.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-11.

15. A computer program product comprising a computer program, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.