Method, device, equipment and storage medium for determining inference data

By screening and building a multimedia reasoning database, the problem of scarcity of multimodal data in large models is solved, efficient use of web data is achieved, and the reasoning ability and adaptability of the model are improved.

CN119312914BActive Publication Date: 2025-09-23BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411327972.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-23
Publication Date
2025-09-23
Estimated Expiration
2044-09-23

AI Technical Summary

Technical Problem

In multimodal applications of large models, high-quality multimodal reasoning data is scarce, and the quality of existing web page data is uneven, resulting in inefficient utilization.

Method used

By screening the first web page that meets the page quality requirements and the second web page that meets some data quality requirements, the parsed data is used to build a multimedia reasoning database, including data in multiple modalities such as text, image and voice.

Benefits of technology

Efficiently mine high-quality multimodal reasoning data from massive web pages, improving the data richness and reasoning capabilities of model training and adapting to various languages ​​and fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119312914B_ABST
    Figure CN119312914B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, apparatus, device, and storage medium for determining inference data, relating to the fields of data processing technology, particularly artificial intelligence, big data, and large models. A specific implementation scheme comprises: obtaining a first set of web pages and a second set of web pages, wherein the first set of web pages includes at least one first web page whose overall data quality meets page quality requirements, and the second set of web pages includes at least one second web page whose data quality of at least part of the web page content meets data requirements; and obtaining a multimedia inference database for model training using at least first parsed data of the first web page and second parsed data of the second web page.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to technical fields such as artificial intelligence, big data, and big models. Background Art

[0002] In large-scale multimodal applications, high-quality multimodal reasoning data is severely scarce, limiting the reasoning capabilities of large models. While existing web pages contain rich multimodal data, the quality of this data varies widely, hindering efficient utilization of web data. For example, cleaning, processing, and mining web data remain key challenges. Summary of the Invention

[0003] The present disclosure provides a method, apparatus, device, and storage medium for determining inference data.

[0004] According to one aspect of the present disclosure, a method for determining inference data is provided, comprising:

[0005] Acquire a first web page set and a second web page set, wherein the first web page set includes at least one first web page whose overall data quality meets the page quality requirements, and the second web page set includes at least one second web page whose data quality of at least part of the web page content meets the data requirements;

[0006] A multimedia inference database for model training is obtained by at least utilizing the first parsed data of the first webpage and the second parsed data of the second webpage.

[0007] According to another aspect of the present disclosure, there is provided a device for determining inference data, comprising:

[0008] a data preprocessing unit configured to obtain a first web page set and a second web page set, wherein the first web page set includes at least one first web page whose overall data quality meets the page quality requirements, and the second web page set includes at least one second web page whose data quality of at least part of the web page content meets the data requirements;

[0009] The database construction unit is used to obtain a multimedia reasoning database for model training by at least using the first parsed data of the first web page and the second parsed data of the second web page.

[0010] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0011] at least one processor; and

[0012] a memory communicatively connected to the at least one processor; wherein,

[0013] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method in the embodiments of the present disclosure.

[0014] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.

[0015] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements any one of the methods according to the embodiments of the present disclosure.

[0016] The disclosed solution can utilize different screening requirements (such as page quality requirements and data requirements) to determine the required web pages (such as the first web page and the second web page), and then use the first parsed data of the first web page and the second parsed data of the second web page to construct inference data that can be used for model training (such as a multimedia inference database). In this way, high-quality inference data under multimodality can be efficiently mined from a large number of web pages, thereby improving the richness of the inference data and laying the foundation for improving the reasoning ability of the model in the subsequent model training process.

[0017] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0019] Figure 1 This is a schematic flow chart of a method for determining inference data according to an embodiment of the present application. Figure 1 ;

[0020] Figure 2 This is a schematic flow chart of a method for determining inference data according to an embodiment of the present application. Figure 2 ;

[0021] Figure 3 This is a schematic flow chart of a method for determining inference data according to an embodiment of the present application. Figure 3 ;

[0022] Figure 4 is a flowchart of a method for determining inference data according to an embodiment of the present application in a specific example;

[0023] Figure 5is a schematic structural diagram of an apparatus for determining inference data according to an embodiment of the present application;

[0024] Figure 6 It is a block diagram of an electronic device used to implement the method for determining inference data according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0026] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The term "at least one" in this article means any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C, can mean including any one or more elements selected from the set consisting of A, B, and C. The terms "first" and "second" in this article refer to multiple similar technical terms and distinguish them, and do not mean to limit the order or to limit to only two. For example, the first feature and the second feature refer to two categories / two features. The first feature can be one or more, and the second feature can also be one or more.

[0027] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0028] The disclosed solution proposes a method for determining inference data to obtain high-quality inference data.

[0029] Figure 1 This is a schematic flow chart of a method for determining inference data according to an embodiment of the present application. Figure 1 The method may be optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.

[0030] Furthermore, the method includes at least part of the following contents. Figure 1 Shown, including:

[0031] Step S101: Acquire a first web page set and a second web page set.

[0032] Here, the first web page set includes at least one first web page whose overall data quality meets the page quality requirements, and the second web page set includes at least one second web page whose data quality of at least part of the web page content meets the data requirements.

[0033] For example, in one example, the overall data quality of the first web page included in the first web page set meets the page quality requirements, and the overall data quality of the second web page included in the second web page set does not meet the page quality requirements, but the data quality of some page contents meets the data requirements.

[0034] Step S102: obtaining a multimedia inference database for model training by at least utilizing the first parsed data of the first webpage and the second parsed data of the second webpage.

[0035] That is, in this example, the first parsed data of each first web page in the first web page set and the second parsed data of each second web page in the second web page set can be used to obtain a multimedia reasoning database that can be used for model training.

[0036] In this way, the disclosed solution can use different screening requirements (such as page quality requirements and data requirements) to determine the required web pages (such as the first web page and the second web page), and then use the first parsed data of the first web page and the second parsed data of the second web page to construct inference data that can be used for model training (such as a multimedia inference database). In this way, high-quality inference data under multimodality can be efficiently mined from a large number of web pages, thereby improving the richness of the inference data and laying the foundation for improving the reasoning ability of the model in the subsequent model training process.

[0037] Furthermore, since the web pages involved in the disclosed solution (such as the first web page and the second web page) can be web pages in any field and expressed in any language, the disclosed solution can obtain reasoning data in different languages ​​(such as Chinese, English, etc.) and different modalities from web pages in different fields. At this time, if the obtained reasoning data is used for model training, the subsequently trained model can have the reasoning ability to adapt to various languages ​​and various fields. Therefore, the disclosed solution has strong applicability and practicality.

[0038] It should be noted that, in a specific example, the disclosed solution can also construct a multimedia inference database for model training using the first parsed data of at least one first web page when a first set of web pages is obtained, or construct a multimedia inference database for model training using the second parsed data of at least one second web page when a second set of web pages is obtained. The disclosed solution does not impose any specific restrictions on this.

[0039] Furthermore, in one example, the multimedia reasoning database may contain reasoning data in multiple modalities (e.g., images, text, and / or voice data formats). In this case, it may also be specifically referred to as a multimodal reasoning database. For example, the multimedia database includes, but is not limited to, at least one of the following: text data, image data, audio data, video data, etc., wherein the text data may also be text data in different languages, such as Chinese text data or English text data. This disclosure is not limited to this.

[0040] Figure 2 This is a schematic flow chart of a method for determining inference data according to an embodiment of the present application. Figure 2 The method can be optionally applied to electronic devices, such as personal computers, servers, server clusters and other electronic devices. It is understood that the above Figure 1 The relevant contents of the method shown can also be applied to this example, and this example will not elaborate on the relevant contents.

[0041] Furthermore, the method includes at least part of the following contents. Figure 2 Shown, including:

[0042] Step S201: Obtain target parsing data of each of a plurality of preset web pages.

[0043] Step S202: Based on the target parsed data of the preset webpage, the overall data quality of the preset webpage is evaluated to obtain a page evaluation result of the preset webpage.

[0044] Step S203: taking a preset web page whose page evaluation result meets the page quality requirement as a first web page to obtain a first web page set.

[0045] Here, the first parsed data of the first webpage is the target parsed data. Further, the page evaluation result may be specifically a page evaluation score. In this case, the page quality requirement may be specifically that the page evaluation score is greater than or equal to a first threshold.

[0046] Step S204: Acquire a second web page set.

[0047] Here, the second web page set includes at least one second web page of which the data quality of at least a portion of the web page content meets the data requirements.

[0048] It is understandable that in actual applications, the step of obtaining the second web page set (i.e., step S204) can be interchanged with any step of obtaining the first web page set (i.e., steps S201 to S203), or the two can be performed simultaneously, and the present disclosure does not limit this.

[0049] Step S205: Obtain a multimedia inference database for model training by at least utilizing the first parsed data of the first webpage and the second parsed data of the second webpage.

[0050] That is to say, in this example, first, based on the target parsing data of the preset web page obtained, the overall data quality of the preset web page is evaluated to obtain the page evaluation result of the preset web page (such as the page evaluation score) to obtain the page evaluation result of each preset web page; secondly, from all the preset web pages, the preset web page whose page evaluation result meets the page quality requirements is selected as the first web page. For example, in one example, the preset web page whose page evaluation score is greater than or equal to the first threshold is selected as the first web page to obtain the first web page set. At this time, the first parsing data of the first web page in the first web page set can be specifically the target parsing data; finally, the first parsing data of at least one first web page and the second parsing data of the obtained second web page are used to construct a multimedia reasoning database.

[0051] In this way, the disclosed solution provides a method for obtaining a first set of web pages, which can quickly and efficiently obtain the first web page (i.e., the first set of web pages) that meets the page quality requirements. In this way, it provides data support for the subsequent rapid construction of a multimedia reasoning database, and thus lays the foundation for subsequent training to obtain an inference model with stronger reasoning capabilities.

[0052] Furthermore, in a specific example, the following method can be used to quickly obtain a page evaluation result. Specifically, based on the target parsed data of the preset web page, the overall data quality of the preset web page is evaluated to obtain the page evaluation result of the preset web page (for example, step S202), which specifically includes:

[0053] The target parsed data of the preset web page is input into the data quality model to obtain the page evaluation result of the preset web page.

[0054] Here, the data quality model is used to predict the data quality of the entire web page and output an evaluation result.

[0055] For example, in one example, the data quality model can output a page evaluation score. In this case, the first web page that meets the page quality requirements can be quickly screened based on the page evaluation score output by the data quality model.

[0056] Alternatively, in one example, the data quality model can be specifically a classification model. In this case, the output evaluation results can be specifically classification results. For example, classification results of category 1 or category 2 can be obtained, where category 1 can indicate that the page evaluation results meet the page quality requirements, and category 2 can indicate that the page evaluation results do not meet the page quality requirements.

[0057] It should be noted that the disclosed solution does not impose any specific restrictions on the data quality model, and the model can be selected or trained based on actual needs.

[0058] In this way, the disclosed solution can use the data quality model to evaluate the overall data quality of each preset web page, and thus quickly and accurately obtain the page evaluation results of each preset web page. This provides strong support for obtaining the first web page and the second web page, and lays the foundation for the subsequent construction of a high-quality multimedia reasoning database.

[0059] Furthermore, in a specific example, the target parsed data of the preset web page can be obtained in the following manner; specifically, the above-mentioned obtaining of the target parsed data of each preset web page in the plurality of preset web pages (e.g., step S201) specifically includes:

[0060] Step S201 - 1 : Obtaining initial webpage data of each of the plurality of preset webpages.

[0061] For example, in one example, the initial page data may be specifically HyperText Markup Language (HTML) data.

[0062] Step S201 - 2 : parsing the initial web page data of the preset web pages based on a parsing method that matches the initial web page data to obtain target parsed data of each preset web page.

[0063] For example, in one example, the initial web page data may contain a specified tag (for example, "table", "audio", "code" or "video", etc.). At this time, the parsing method corresponding to the initial web page data can be determined based on the specified tag. For example, if the initial web page data contains the specified tag "table", the parsing method corresponding to the specified tag "table" is used to parse the initial web page data. For example, the data where the specified tag "table" is located can be extracted to obtain the target parsed data of the preset web page.

[0064] Alternatively, in another example, the initial web page data may not include the specified tag. In this case, the initial web page data may be parsed based on a general parsing method to obtain target parsed data.

[0065] In this way, the disclosed solution provides a refined solution for obtaining target parsed data, which can provide data support for subsequent scientific and reasonable evaluation of the overall data quality of preset web pages, thereby laying the foundation for building a high-quality multimedia reasoning database.

[0066] Figure 3 This is a schematic flow chart of a method for determining inference data according to an embodiment of the present application. Figure 3 The method can be optionally applied to electronic devices, such as personal computers, servers, server clusters and other electronic devices. It is understood that the above Figure 1 and Figure 2 The relevant contents of the method shown can also be applied to this example, and this example will not elaborate on the relevant contents.

[0067] Furthermore, the method includes at least part of the following contents. Figure 3 Shown, including:

[0068] Step S301: Obtain target parsing data of each of a plurality of preset web pages.

[0069] Step S302: Based on the target parsed data of the preset webpage, the overall data quality of the preset webpage is evaluated to obtain a page evaluation result of the preset webpage.

[0070] It should be noted that the preset web page parsed using the special parsing method does not need to undergo overall data quality assessment of the web page. The preset web page can be directly used as the third web page to obtain a third web page set. At this time, the third parsed data of the third web page is the target parsed data of the preset web page.

[0071] Step S303: Determine whether the page evaluation result of the preset web page meets the page quality requirement. If so, proceed to step S304; otherwise, proceed to step S305.

[0072] Step S304: taking a preset web page whose page evaluation result meets the page quality requirement as a first web page to obtain a first web page set.

[0073] Here, the first parsed data of the first webpage is the target parsed data of the preset webpage.

[0074] Step S305: Determine whether there is target data meeting the data requirements in the target parsed data of the preset web page whose page evaluation result does not meet the page quality requirement. If so, proceed to step S306; otherwise, proceed to step S307.

[0075] For example, in one example, it can be specifically determined whether the content such as a paragraph, an image link, or a text fragment in the target parsed data of a preset web page meets the data requirements.

[0076] Step S306: When it is determined that target data meeting the data requirement exists, the preset web page containing the target data is used as the second web page to obtain a second web page set.

[0077] That is, in this example, if the page evaluation result of the preset webpage is determined not to meet the page quality requirements, the target parsed data of the preset webpage can be further determined to determine whether there is target data that meets the data requirements, such as whether there is a paragraph, image link, text fragment, etc. that meets the data requirements. If so, the preset webpage that contains the target data that meets the data requirements can be used as the second webpage.

[0078] Furthermore, in this example, the second parsed data of the second webpage may specifically be target data that meets the data requirement in the target parsed data of the second webpage.

[0079] Step S307: Abandon the preset web page.

[0080] In this way, by using the above steps S302 to S307 , the evaluation and classification of all the preset web pages can be completed, and the first web page set and the second web page set can be obtained, and then the process proceeds to step S308 .

[0081] Step S308: Obtain a multimedia inference database for model training using at least the first parsed data of the first webpage and the second parsed data of the second webpage.

[0082] For example, in one example, first, based on the target parsed data of the preset webpage obtained, the overall data quality of the preset webpage is evaluated to obtain a page evaluation score for each preset webpage; second, it is determined whether the page evaluation score of the preset webpage is greater than or equal to a first threshold. If the page evaluation score of the preset webpage is greater than or equal to the first threshold, the preset webpage is used as the first webpage to obtain a first webpage set; otherwise, that is, when the page evaluation score of the preset webpage is less than the first threshold, it is determined whether the target parsed data of the preset webpage contains target data that meets the data requirements. If so, the preset webpage is used as the second webpage to obtain a second webpage set; otherwise, the preset webpage is discarded. At this point, the evaluation and classification of all preset webpages are completed; finally, the first parsed data of at least one first webpage and the second parsed data of at least one second webpage can be used to construct a multimedia reasoning database.

[0083] In this way, the disclosed solution provides a specific solution that can evaluate and classify multiple preset web pages. In this way, a first set of web pages and a second set of web pages can be efficiently obtained, and then the first parsed data of the first web page and the second parsed data of the second web page can be used to quickly construct a multimedia reasoning database. In this way, high-quality reasoning data can be effectively obtained, thereby laying the foundation for improving the reasoning ability of the model in the subsequent model training process.

[0084] Furthermore, in a specific example, the multimedia inference database can be obtained in the following manner; specifically, the above-described method of using at least the first parsed data of the first webpage and the second parsed data of the second webpage to obtain the multimedia inference database for model training (e.g., step S102, step S205, step S308) specifically includes:

[0085] Step S308 - 1 : Acquire multimedia data corresponding to first parsed data of a first webpage, and acquire multimedia data corresponding to second parsed data of a second webpage, to obtain a multimedia data set.

[0086] For example, in one example, multimedia data download is performed on the parsed data of each web page to obtain the multimedia data corresponding to each web page, and then obtain a multimedia data set.

[0087] Step S308 - 2 : Label each multimedia data in the multimedia data set to obtain multimedia data with labeling information.

[0088] Here, the multimedia data to be annotated with information may also be referred to as inference data.

[0089] Step S308 - 3 : obtaining a multimedia reasoning database for model training based at least on the multimedia data with annotation information.

[0090] It should be noted that, in one example, multimedia data with annotated information can constitute a multimedia inference database, while multimedia data that has not been annotated, or has failed to be annotated, can constitute general data. Alternatively, if the annotated information includes a preset label and a score (e.g., the probability of the data corresponding to the preset label), multimedia data with a score greater than or equal to the preset score can be used to construct a multimedia inference database; multimedia data with a score less than the preset score can be assigned to a general database. In this way, the disclosed solution can construct a multimedia inference database for model training on the one hand, and a general database on the other.

[0091] In this way, the disclosed solution can annotate the obtained multimedia data, and then use the multimedia data with annotated information to construct a multimedia reasoning database. In this way, the quality and diversity of the multimedia reasoning data are improved, laying the foundation for improving the reasoning ability of the model in the subsequent model training process.

[0092] Furthermore, in a specific example, the tagging information of each multimedia data may be obtained in the following manner. Specifically, the tagging of each multimedia data in the multimedia data set to obtain multimedia data with the tagging information (e.g., step S308-2) described above may include at least one of the following:

[0093] Annotation method 1: input the image data in the multimedia data into the image reasoning annotation model to obtain the annotation information of the image data in the multimedia data.

[0094] Annotation method 2: Input the text data in the multimedia data into the text reasoning annotation model to obtain the annotation information of the text data in the multimedia data.

[0095] In this way, the disclosed solution can use different annotation models according to different data types in the multimedia data to quickly annotate the corresponding data in the multimedia data, so as to quickly obtain multimedia data with annotation information, thus laying the foundation for the subsequent construction of a high-quality multimedia reasoning database.

[0096] In a specific example of the disclosed solution, after constructing the multimedia database, the process further includes obtaining training samples for training the aforementioned annotation model (e.g., an image reasoning annotation model or a text reasoning annotation model) (e.g., step S309), which may specifically include:

[0097] Step S309 - 1 : Based on the first webpage and / or the second webpage obtained from the multimedia reasoning database, a plurality of initial websites are obtained.

[0098] For example, in one example, a plurality of initial websites in the form of “XXX.com” may be obtained based on the first web page and / or the second web page used to construct the multimedia reasoning database.

[0099] Here, it should be noted that the website described in this example may refer to a website, including a set of web pages provided by the same domain and maintained by a single organization. In other words, a website may include multiple web pages that can be interconnected via hyperlinks.

[0100] Step S309 - 2 : Based on the multiple initial websites, determine at least one target page that meets the page requirements.

[0101] That is, in this example, at least one target page that meets the page requirements can be determined based on the multiple web pages contained in the initial website. For example, in one example, the web pages that can be clicked and opened in the initial website can be annotated using the macro model, and the web pages annotated with "inference" information are used as target pages (also called target web pages). In this way, multiple target pages can be obtained.

[0102] It should be noted that the large model can mark pages that contain sample data that can be used for model training as "inference" pages.

[0103] Step S309 - 3 : Using the at least one target page, obtain image samples for training the image reasoning annotation model and / or text samples for training the text reasoning annotation model.

[0104] In other words, in one example, first, based on the first and / or second web pages obtained from the multimedia inference database, multiple initial websites are statistically obtained. Second, for each initial website, if the web pages that can be clicked and opened under the initial website can be annotated using the large model, and the web pages annotated with "inference" information are used as target pages, at least one target page can be obtained. Finally, from each obtained target page, at least one of the following samples is obtained: an image sample or a text sample. This provides data support for subsequently improving the annotation accuracy of the inference annotation model.

[0105] In this way, the disclosed solution provides a method for efficiently obtaining training samples for training image reasoning annotation models and / or text reasoning annotation models, thereby providing data support for subsequent improvement of the annotation quality of the reasoning annotation model, and laying the foundation for improving the richness and data quality of reasoning data in multimedia reasoning databases.

[0106] Furthermore, in a specific example, after obtaining the image sample and / or the text sample, the method further includes:

[0107] (1) Training the model used in labeling method 1 (i.e., the image reasoning labeling model). Specifically, at least using the image samples obtained for training the image reasoning labeling model, the image reasoning labeling model is trained to obtain the image reasoning labeling model after training.

[0108] And / or, (2) training the model (text inference annotation model) used in annotation method 2, specifically, performing model training on the text inference annotation model using at least the text samples obtained for training the text inference annotation model to obtain the text inference annotation model after training.

[0109] For example, in one example, the trained image reasoning annotation model and text reasoning annotation model can further annotate unlabeled multimedia data, such as multimedia data in a general database, to further enrich the data volume of the multimedia reasoning database.

[0110] In this way, the disclosed solution can use the obtained image samples to train the image reasoning annotation model, and / or use the obtained text samples to train the text reasoning annotation model. In this way, the annotation accuracy of the reasoning annotation model is effectively improved, laying the foundation for the subsequent improvement of the data volume, data quality and data richness of the constructed multimedia reasoning database.

[0111] Furthermore, in a specific example, at least one target page may be obtained in the following manner; specifically, the above-described determination of at least one target page that meets the page requirements based on the multiple initial websites (e.g., step S309-2) specifically includes:

[0112] Step S309-2-1: Determine at least one target network site that meets preset site requirements from the multiple initial network sites.

[0113] For example, in one example, an initial website whose number of web pages that can be clicked and opened exceeds a second threshold can be selected as the target website. Alternatively, in another example, an initial website whose number of web pages that can be clicked and opened exceeds a second threshold and whose number of web pages marked as "inference" among the clicked and opened web pages exceeds a third threshold can be selected as the target website. In other words, an initial website that meets the above two conditions can be selected as the target website.

[0114] It should be noted that the above is only an exemplary description. In actual applications, the preset site requirements may be determined based on actual scenario needs, and the present disclosure does not limit this.

[0115] Step S309-2-2: Select at least one target page that meets the page requirement from the at least one target website.

[0116] For instance, in one example, when the number of web pages that can be clicked to open under the initial website exceeds 100, and the number of web pages that are "inference" web pages exceeds 10%, the initial website can be used as the target website to obtain at least one target website; and then 30% of the web pages of each target website are sampled (hereinafter referred to as sampled web pages), and the large model is used to determine whether each sampled web page belongs to an "inference" web page, and the sampled web pages that belong to the "inference" web page are used as target pages to obtain at least one target page.

[0117] In this way, the disclosed solution provides a refined solution for obtaining the target page, which is simple, efficient, and highly interpretable. This lays the foundation for the subsequent efficient acquisition of samples that can be used to train image reasoning annotation models and text reasoning annotation models, and further lays the foundation for improving the annotation accuracy of image reasoning annotation models and text reasoning annotation models, and for improving the data volume, data quality, and data richness of the constructed multimedia reasoning database.

[0118] The following is a detailed description of the disclosed solution with reference to specific examples. The disclosed solution provides a solution for mining inference data. Specifically, the disclosed solution can determine the mining direction on reasoning and define an inference label system according to actual needs. For example, in one example, the inference label system may include explanations, steps, authoritative arguments, mathematics, analogy arguments, etc. In this way, the multimedia data obtained by subsequent mining can be labeled based on the defined inference label system to obtain the required inference data.

[0119] Furthermore, the disclosed solution can also produce target domain data (corresponding to the above-mentioned multimedia reasoning database) and general data (compared to multimedia reasoning data, the database composed of general data can be called a general database) through processes such as data analysis, data cleaning, and model annotation on web pages; and, the disclosed solution can also further improve the data volume, data quality, and data richness of the produced target domain data through model training, correction, and recycling. For example, the disclosed solution can also obtain seed data (corresponding to the target page mentioned above) for training or optimizing the multimodal reasoning annotation model to obtain a text reasoning annotation model and an image reasoning annotation model, thereby improving the quantity, data quality, and richness of the reasoning data.

[0120] Specifically, if Figure 4 As shown, the steps of the disclosed solution include:

[0121] Step S401: Determine whether the HTML data of the webpage (corresponding to the initial webpage data) contains a specified tag (such as table, audio, code, video, etc.) to determine a parsing method that matches the HTML data.

[0122] Here, in one example, if it is determined that the HTML data of the web page contains a specified tag, the HTML data of the web page can be parsed using a preset special parsing method to obtain target parsed data and directly enter step S404; otherwise, a general parsing method is used to parse the HTML data of the web page to obtain target parsed data and enter step S402.

[0123] Step S402: Input the target parsed data of each web page into the data quality model to classify the target parsed data and give a quality classification label. For example, if the quality classification label of the web page indicates that it meets the high quality requirements, then the web page is used as the first web page to obtain the first web page set. At this time, the parsed data of the first web page can be specifically the target parsed data. Otherwise (that is, the quality classification label of the web page indicates that it does not meet the high quality requirements), continue to determine whether the target parsed data of the web page contains target data that meets the data requirements (such as legal image links, legal text links). If the target data exists, then the web page is used as the second web page. At this time, the target parsed data of the second web page is filtered, and the filtered target parsed data, for example, the target data that meets the data requirements is directly used as the parsed data of the second web page. Otherwise (that is, there is no target data), the preset web page is discarded, such as being placed in a trash can.

[0124] Step S403: Download multimedia data (also called multimedia materials) from the parsed data of the first web page and the parsed data of the second web page to obtain multimedia data (such as pictures, videos, audios, etc.) corresponding to each web page, and then obtain a multimedia corpus (corresponding to the above multimedia data set).

[0125] Step S404: Combine the image reasoning annotation model and the text reasoning annotation model to annotate the multimedia data in the multimedia corpus to obtain multimedia data with annotation information to build a multimedia reasoning library (that is, corresponding to the above multimedia reasoning database). In addition, a multimedia general library (that is, corresponding to the above general database) is also obtained.

[0126] Here, in one example, an image reasoning annotation model is used to annotate image data in multimedia data, and annotation information of the image data (e.g., including preset reasoning labels and scores) is obtained; a text reasoning annotation model is used to annotate text data in multimedia data, and annotation information of the text data (e.g., including preset reasoning labels and scores) is obtained; further, image data and text data with scores greater than or equal to a preset score are used to construct a multimedia reasoning library, and image data and text data with scores less than a preset score are used to construct a multimedia general library.

[0127] Furthermore, it is also possible to construct a text inference database using only text data with scores greater than or equal to a preset score, and to construct a text general database using text data with scores less than a preset score. In this way, multiple databases can be constructed.

[0128] Step S405: Based on the first web page and the second web page of the multimedia reasoning library, a plurality of initial network sites are obtained by counting, and when it is determined that the total number of web pages under the initial network site exceeds 100 and more than 10% of them are determined to be reasoning web pages, the initial network site is used as a suspected reasoning site (corresponding to the above target network site) to obtain at least one suspected reasoning site.

[0129] Step S406: Sample 30% of the web pages for each suspected inference site (i.e., corresponding to the sampled web pages above), use the large model to annotate these web pages, and use the web pages labeled "inference" to obtain text samples for training the text inference annotation model, and obtain image samples for training the image inference annotation model.

[0130] Step S407: Using the obtained text samples, the text inference annotation model in step S404 is trained, and using the obtained image samples, the image inference annotation model in step S404 is trained, thereby further enhancing the diversity of the model.

[0131] Step S408: loop through steps S404 to S407 to obtain a multimedia reasoning library consisting of a large amount of high-quality multimodal reasoning data.

[0132] In this way, the disclosed solution can combine text reasoning and annotation models with image reasoning and annotation models to mine high-quality multimodal reasoning corpus. Compared with existing solutions, it enriches data diversity and lays the foundation for improving the model's reasoning ability during subsequent model training.

[0133] The disclosed solution also provides a device for determining inference data, such as Figure 5 Shown, including:

[0134] A data preprocessing unit 501 is configured to obtain a first web page set and a second web page set, wherein the first web page set includes at least one first web page whose overall data quality meets page quality requirements, and the second web page set includes at least one second web page whose data quality of at least part of the web page content meets data requirements;

[0135] The database construction unit 502 is configured to obtain a multimedia inference database for model training using at least the first parsed data of the first web page and the second parsed data of the second web page.

[0136] In a specific example of the disclosed solution, the data preprocessing unit is further configured to:

[0137] Obtaining target parsed data for each of a plurality of preset web pages;

[0138] Based on the target parsed data of the preset web page, the overall data quality of the preset web page is evaluated to obtain a page evaluation result of the preset web page;

[0139] A preset web page whose page evaluation result meets the page quality requirement is used as a first web page to obtain the first web page set; wherein the first parsed data of the first web page is the target parsed data.

[0140] In a specific example of the disclosed solution, the data preprocessing unit is specifically configured to:

[0141] The target parsed data of the preset web page is input into the data quality model to obtain a page evaluation result of the preset web page, wherein the data quality model is used to predict the data quality of the entire web page and output the evaluation result.

[0142] In a specific example of the disclosed solution, the data preprocessing unit is specifically configured to:

[0143] Obtaining initial webpage data of each preset webpage among the plurality of preset webpages;

[0144] Based on the parsing method that matches the initial web page data, the initial web page data of the preset web page is parsed to obtain target parsed data of each preset web page.

[0145] In a specific example of the disclosed solution, the data preprocessing unit is further configured to:

[0146] determining whether there is target data meeting the data requirements in the target parsed data of the preset web page whose page evaluation result does not meet the page quality requirements;

[0147] When it is determined that target data that meets the data requirements exists, the preset web page containing the target data is used as the second web page to obtain the second web page set; wherein the second parsed data of the second web page is: the target data that meets the data requirements in the target parsed data of the second web page.

[0148] In a specific example of the present disclosure, the database construction unit is specifically configured to:

[0149] Acquire multimedia data corresponding to first parsed data of the first webpage, and acquire multimedia data corresponding to second parsed data of the second webpage, to obtain a multimedia data set;

[0150] Marking each multimedia data in the multimedia data set to obtain multimedia data with marking information;

[0151] At least based on the multimedia data with labeled information, a multimedia reasoning database for model training is obtained.

[0152] In a specific example of the disclosed solution, the database construction unit specifically performs at least one of the following:

[0153] Inputting image data in the multimedia data into the image reasoning annotation model to obtain annotation information of the image data in the multimedia data;

[0154] The text data in the multimedia data is input into the text inference annotation model to obtain the annotation information of the text data in the multimedia data.

[0155] In a specific example of the disclosed solution, the system further includes a site processing unit, wherein the site processing unit is configured to:

[0156] Based on the first web page and / or the second web page of the multimedia reasoning database, a plurality of initial websites are obtained;

[0157] Determining at least one target page that meets the page requirements based on the multiple initial websites;

[0158] The at least one target page is used to obtain image samples for training the image reasoning annotation model and / or text samples for training the text reasoning annotation model.

[0159] In a specific example of the disclosed solution, the system further includes a training unit, wherein the training unit is configured to:

[0160] Performing model training on the image reasoning and annotation model using at least the obtained image samples for training the image reasoning and annotation model to obtain the trained image reasoning and annotation model;

[0161] and / or,

[0162] The text reasoning annotation model is at least trained using the obtained text samples for training the text reasoning annotation model to obtain the text reasoning annotation model after training.

[0163] In a specific example of the disclosed solution, the site processing unit is specifically configured to:

[0164] Determining at least one target network site that meets preset site requirements from the multiple initial network sites;

[0165] At least one target page that meets the page requirement is selected from the at least one target website.

[0166] For the description of specific functions and examples of each unit of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0167] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0168] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0169] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0170] like Figure 6 As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0171] Various components in device 600 are connected to I / O interface 605, including an input unit 606, such as a keyboard, mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, optical disk, etc.; and a communication unit 609, such as a network card, modem, wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0172] The computing unit 601 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as the method for determining inference data. For example, in some embodiments, the method for determining inference data can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method for determining inference data described above can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to perform the method for determining inference data by any other appropriate means (e.g., by means of firmware).

[0173] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0174] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0175] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0176] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0177] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0178] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0179] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0180] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for determining inference data, comprising: Acquire a first web page set and a second web page set, wherein the first web page set includes at least one first web page whose overall data quality meets the page quality requirements, and the second web page set includes at least one second web page whose partial web page content data quality meets the data requirements; Acquire multimedia data corresponding to first parsed data of the first webpage, and acquire multimedia data corresponding to second parsed data of the second webpage, to obtain a multimedia data set; Marking each multimedia data in the multimedia data set to obtain multimedia data with marking information; At least based on the multimedia data with labeled information, a multimedia reasoning database for model training is obtained; Among them, also include: Based on the first web page and / or the second web page of the multimedia reasoning database, a plurality of initial websites are obtained; Determining at least one target page that meets the page requirements based on the multiple initial websites; The at least one target page is used to obtain image samples for training an image reasoning annotation model and / or text samples for training a text reasoning annotation model.

2. The method according to claim 1, further comprising: Obtaining target parsed data for each of a plurality of preset web pages; Based on the target parsed data of the preset web page, the overall data quality of the preset web page is evaluated to obtain a page evaluation result of the preset web page; A preset web page whose page evaluation result meets the page quality requirement is used as a first web page to obtain the first web page set; wherein the first parsed data of the first web page is the target parsed data.

3. The method according to claim 2, wherein: The target parsing data of the preset web page is used to evaluate the overall data quality of the preset web page to obtain a page evaluation result of the preset web page, including: The target parsed data of the preset web page is input into the data quality model to obtain a page evaluation result of the preset web page, wherein the data quality model is used to predict the data quality of the entire web page and output the evaluation result.

4. The method according to claim 2, wherein: The step of obtaining target parsed data of each of the plurality of preset web pages includes: Obtaining initial webpage data of each preset webpage among the plurality of preset webpages; Based on the parsing method that matches the initial web page data, the initial web page data of the preset web page is parsed to obtain target parsed data of each preset web page.

5. The method according to claim 2, further comprising: determining whether there is target data meeting the data requirements in the target parsed data of the preset web page whose page evaluation result does not meet the page quality requirements; When it is determined that target data that meets the data requirements exists, the preset web page containing the target data is used as the second web page to obtain the second web page set; wherein the second parsed data of the second web page is: the target data that meets the data requirements in the target parsed data of the second web page.

6. The method according to any one of claims 1 to 5, wherein: The step of labeling each multimedia data in the multimedia data set to obtain multimedia data with labeling information includes at least one of the following: Inputting image data in the multimedia data into the image reasoning annotation model to obtain annotation information of the image data in the multimedia data; The text data in the multimedia data is input into the text inference annotation model to obtain the annotation information of the text data in the multimedia data.

7. The method according to any one of claims 1 to 5, further comprising: Performing model training on the image reasoning and annotation model using at least the obtained image samples for training the image reasoning and annotation model to obtain the trained image reasoning and annotation model; and / or, The text reasoning annotation model is at least trained using the obtained text samples for training the text reasoning annotation model to obtain the text reasoning annotation model after training.

8. The method according to any one of claims 1 to 5, wherein: The step of determining at least one target page that meets the page requirements based on the multiple initial websites includes: Determining at least one target network site that meets preset site requirements from the multiple initial network sites; At least one target page that meets the page requirement is selected from the at least one target website.

9. A device for determining inference data, comprising: a data preprocessing unit configured to obtain a first web page set and a second web page set, wherein the first web page set includes at least one first web page whose overall data quality meets the page quality requirements, and the second web page set includes at least one second web page whose partial web page content data quality meets the data requirements; A database construction unit is configured to obtain multimedia data corresponding to first parsed data of a first webpage and multimedia data corresponding to second parsed data of a second webpage to obtain a multimedia data set; annotate each multimedia data in the multimedia data set to obtain multimedia data with annotated information; and obtain a multimedia inference database for model training based at least on the multimedia data with annotated information; The device further includes a site processing unit, wherein: The site processing unit is used to obtain multiple initial network sites based on the first web page and / or the second web page obtained from the multimedia reasoning database; determine at least one target page that meets the page requirements based on the multiple initial network sites; and use the at least one target page to obtain image samples for training the image reasoning annotation model and / or text samples for training the text reasoning annotation model.

10. The device according to claim 9, wherein The data preprocessing unit is further used for: Obtaining target parsed data for each of a plurality of preset web pages; Based on the target parsed data of the preset web page, the overall data quality of the preset web page is evaluated to obtain a page evaluation result of the preset web page; A preset web page whose page evaluation result meets the page quality requirement is used as a first web page to obtain the first web page set; wherein the first parsed data of the first web page is the target parsed data.

11. The device according to claim 10, wherein The data preprocessing unit is specifically used to: The target parsed data of the preset web page is input into the data quality model to obtain a page evaluation result of the preset web page, wherein the data quality model is used to predict the data quality of the entire web page and output the evaluation result.

12. The device according to claim 10, wherein The data preprocessing unit is specifically used to: Obtaining initial webpage data of each preset webpage among the plurality of preset webpages; Based on the parsing method that matches the initial web page data, the initial web page data of the preset web page is parsed to obtain target parsed data of each preset web page.

13. The apparatus according to claim 10, wherein the data preprocessing unit is further configured to: determining whether there is target data meeting the data requirements in the target parsed data of the preset web page whose page evaluation result does not meet the page quality requirements; When it is determined that target data that meets the data requirements exists, the preset web page containing the target data is used as the second web page to obtain the second web page set; wherein, The second parsed data of the second webpage is: target data that meets the data requirement in the target parsed data of the second webpage.

14. The device according to any one of claims 9 to 13, wherein: The database construction unit specifically performs at least one of the following: Inputting image data in the multimedia data into the image reasoning annotation model to obtain annotation information of the image data in the multimedia data; The text data in the multimedia data is input into the text inference annotation model to obtain the annotation information of the text data in the multimedia data.

15. The apparatus according to any one of claims 9 to 13, further comprising: Training unit; wherein the training unit is used to: Performing model training on the image reasoning and annotation model using at least the obtained image samples for training the image reasoning and annotation model to obtain the trained image reasoning and annotation model; and / or, The text reasoning annotation model is at least trained using the obtained text samples for training the text reasoning annotation model to obtain the text reasoning annotation model after training.

16. The device according to any one of claims 9 to 13, wherein: The site processing unit is specifically configured to: Determining at least one target network site that meets preset site requirements from the multiple initial network sites; At least one target page that meets the page requirement is selected from the at least one target website.

17. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 8.

19. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Data processing method and device, equipment, storage medium and product

    CN118014086A

  • Method and device for constructing data set for model training

    CN118155016A