A web page parsing method and system based on artificial intelligence
By building multimodal models and using adaptive enhancement algorithms, the problem of insufficient adaptability of data extraction in the prior art is solved, and the ability to efficiently extract key information in multimodal data is realized.
Patent Information
- Application Number
- CN202510250754.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-03-04
AI Technical Summary
The prior art lacks the ability to adapt to different application scenarios when extracting key information from multi-source heterogeneous data, resulting in limited adaptability of the model under new tasks or specific needs.
Using an artificial intelligence-based web page analysis method, by obtaining text and image data from multiple data sources, using a two-way long and short-term memory network combined with attention mechanism to build a multimodal model, generating a comprehensive representation vector, and identifying key information through an adaptive enhancement algorithm.
The ability to effectively extract key information from multimodal data is realized, and the model's adaptability to different application scenarios and the quality and efficiency of information processing are improved.
Smart Images

Figure CN119760650B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of multimodal models and information extraction technology, and in particular to a web page parsing method and system based on artificial intelligence. Background Art
[0002] In today's era of information explosion, users are faced with massive amounts of information data, which includes not only text content, but also images, videos and other forms. Especially in the fields of news reporting, social media analysis, e-commerce recommendation systems, medical health record analysis, etc., it is particularly important to quickly and accurately extract key information from multi-source heterogeneous data.
[0003] Most current methods adopt a fixed weight allocation strategy when determining the importance of each feature, without considering the changes in the contribution of each feature in different application scenarios. This limits the model's ability to adapt to new tasks or adjust to meet specific needs. Summary of the invention
[0004] The present application provides a web page parsing method and system based on artificial intelligence to solve the problems of low quality and efficiency of information processing in the prior art.
[0005] In a first aspect, the present application provides a web page parsing method based on artificial intelligence, comprising:
[0006] Acquire associated text data and image data from various data sources;
[0007] Analyzing the text data to capture multi-level semantic information in the text data to obtain serialization information of the text data, and extracting visual features from the image data to identify visual information in the image data to obtain spatial distribution information of the image data, wherein the visual information includes key objects and key scenes;
[0008] A multimodal model is constructed based on a bidirectional long short-term memory network combined with an attention mechanism, wherein the multimodal model can simultaneously process the serialization information of the text data and the spatial distribution information of the image data to determine the complementary association and potential semantic mapping relationship between the text data and the image data;
[0009] Based on the generator obtained by generative adversarial network training, the complementary association and potential semantic mapping relationship output by the multimodal model are fused to generate a comprehensive representation vector, wherein the comprehensive representation vector refers to a high-dimensional vector space expression including multi-level semantic information in the text data, visual information in the image data, complementary associations between the text data and the image data, and potential semantic mapping relationships;
[0010] Using an adaptive enhancement algorithm, dynamically evaluating and adjusting the contribution of each component of the comprehensive representation vector to identify key information, wherein the key information is determined based on multi-level semantic information in the text data, visual information in the image data, and complementary associations and potential semantic mapping relationships between the text data and the image data;
[0011] The key information is extracted from the key information in text form, and representative fragments in the image data are selected to form an information summary. The key points are used to reflect the multi-level semantic information in the text data, and the representative fragments are used to reflect the visual information in the image data.
[0012] Optionally, the use of an adaptive enhancement algorithm to dynamically evaluate and adjust the contribution of each component of the comprehensive representation vector to identify key information includes:
[0013] Initializing a group of weak classifiers using an adaptive boosting algorithm, performing preliminary evaluation processing on each component of the comprehensive representation vector, and obtaining an initial weight of each weak classifier;
[0014] According to the error rate of the current round, the weights of each weak classifier are updated to obtain updated weights;
[0015] Based on the updated weights, each weak classifier is retrained and a new weight coefficient is calculated to obtain a new weight coefficient that reflects the importance of each weak classifier in the final decision;
[0016] Combining all weighted weak classifiers to form a strong classifier, performing global evaluation processing on each component of the comprehensive representation vector, and obtaining an evaluation result of the strong classifier;
[0017] Determining the contribution of each component of the comprehensive representation vector according to the evaluation result of the strong classifier;
[0018] According to the contribution of each component of the comprehensive representation vector, combined with the multi-level semantic information in the text data, the visual information in the image data, the complementary association between the text data and the image data, and the potential semantic mapping relationship, comprehensive analysis and processing are performed to determine the key information.
[0019] Optionally, according to the contribution of each component of the comprehensive representation vector, combined with the multi-level semantic information in the text data, the visual information in the image data, and the complementary association and potential semantic mapping relationship between the text data and the image data, a comprehensive analysis process is performed to determine key information, including:
[0020] Quantifying the contribution of each component of the comprehensive representation vector using the evaluation result of the strong classifier to obtain an importance score for each component;
[0021] Based on the importance score, the components whose contribution meets the set conditions are screened out, and for the screened out components, the hierarchical semantic parsing technology is combined to deeply explore the deep meaning in the text data, and the visual attention mechanism is used to further identify the visual information in the image data, so as to obtain refined multi-level semantic information and visual information;
[0022] Through the cross-modal fusion algorithm, the refined semantic information is integrated with the visual information to generate a multimodal representation;
[0023] Based on the multimodal representation, a graphical model is used to perform structured processing on the refined semantic information and visual information, identify elements with significant consistency between different modalities, and obtain key information.
[0024] Optionally, analyzing the text data to capture multi-level semantic information in the text data to obtain serialization information of the text data, and extracting visual features from the image data to identify visual information in the image data to obtain spatial distribution information of the image data, including:
[0025] Using natural language processing technology, the text data is encoded to obtain a word vector representation;
[0026] Further processing the word vector representation by a sequence modeling method to generate serialization information of the text data;
[0027] Using a convolutional neural network to perform fine-grained visual feature extraction processing on the image data to identify visual information in the image data, wherein the visual information includes key objects and key scenes;
[0028] The specific positions and boundaries of the key objects and key scenes are located and extracted to obtain the spatial distribution information of the image data.
[0029] Optionally, a multimodal model is constructed based on a bidirectional long short-term memory network combined with an attention mechanism, and the multimodal model can simultaneously process the serialization information of the text data and the spatial distribution information of the image data to determine the complementary association and potential semantic mapping relationship between the text data and the image data, including:
[0030] Using a bidirectional long short-term memory network, the serialization information of the text data and the spatial distribution information of the image data are respectively encoded to obtain text features and image features;
[0031] According to the text features and image features output by the bidirectional long short-term memory network, the weight of the bidirectional long short-term memory network output is dynamically adjusted, and the text features and the image features are interacted with each other through a cross attention mechanism, so that the text features can affect the representation of the image features, and the image features can also affect the representation of the text features, so as to establish a complementary association between the text data and the image data;
[0032] A multi-layer perceptron is used to integrate the serialization information of the text data and the spatial distribution information of the image data to form a cross-modal representation to complete the construction of a multimodal model, wherein the cross-modal representation can reflect the potential semantic mapping relationship between the text data and the image data.
[0033] Optionally, based on a generator obtained by generative adversarial network training, the complementary association and potential semantic mapping relationship output by the multimodal model are fused to generate a comprehensive representation vector, including:
[0034] Based on the generator in the generative adversarial network, the complementary associations and potential semantic mapping relationships from the multimodal model are received as input, and the multi-layer neural network structure inside the generator is used to perform nonlinear transformation and deep learning processing on the input complementary associations and potential semantic mapping relationships, so as to achieve effective fusion of the multi-level semantic information and the visual information, and at the same time retain the complementary associations and potential semantic mapping relationships between the text data and the image data;
[0035] The generator outputs a comprehensive representation vector, which not only includes the multi-level semantic information in the text data and the visual information in the image data, but also reflects the complementary relationship and potential semantic mapping relationship between the text data and the image data.
[0036] In a second aspect, the present application provides a web page parsing system based on artificial intelligence, comprising:
[0037] An acquisition module, used for acquiring associated text data and image data from various data sources;
[0038] Analyzing the text data to capture multi-level semantic information in the text data to obtain serialization information of the text data, and extracting visual features from the image data to identify visual information in the image data to obtain spatial distribution information of the image data;
[0039] A construction module, for constructing a multimodal model based on a bidirectional long short-term memory network combined with an attention mechanism, wherein the multimodal model can simultaneously process the serialization information of the text data and the spatial distribution information of the image data to determine the complementary association and potential semantic mapping relationship between the text data and the image data;
[0040] A fusion module, for fusing the complementary associations and potential semantic mapping relationships output by the multimodal model based on a generator trained by a generative adversarial network, and generating a comprehensive representation vector, wherein the comprehensive representation vector refers to a high-dimensional vector space expression including multi-level semantic information in the text data, visual information in the image data, complementary associations between the text data and the image data, and potential semantic mapping relationships;
[0041] A recognition module, for dynamically evaluating and adjusting the contribution of each component of the comprehensive representation vector using an adaptive enhancement algorithm to identify key information, wherein the key information is determined based on multi-level semantic information in the text data, visual information in the image data, and complementary associations and potential semantic mapping relationships between the text data and the image data;
[0042] A generation module is used to extract content highlights in text form from the key information and select representative fragments in the image data to form an information summary. The content highlights are used to reflect the multi-level semantic information in the text data, and the representative fragments are used to reflect the visual information in the image data.
[0043] In a third aspect, the present application provides a computing device comprising a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement an artificial intelligence-based web page parsing method as described in the first aspect above.
[0044] In a fourth aspect, the present application provides a computer storage medium storing a computer program, which, when executed by a computer, implements an artificial intelligence-based web page parsing method as described in the first aspect.
[0045] In the present application, associated text data and image data are obtained from multiple data sources; the text data is analyzed to capture multi-level semantic information in the text data to obtain serialization information of the text data, and visual features are extracted from the image data to identify visual information in the image data to obtain spatial distribution information of the image data, wherein the visual information includes key objects and key scenes; a multimodal model is constructed based on a bidirectional long short-term memory network combined with an attention mechanism, and the multimodal model can simultaneously process the serialization information of the text data and the spatial distribution information of the image data to determine the complementary association and potential semantic mapping relationship between the text data and the image data; based on a generator obtained by generative adversarial network training, the complementary association and potential semantic mapping relationship output by the multimodal model are fused to generate a comprehensive A comprehensive representation vector, wherein the comprehensive representation vector refers to a high-dimensional vector space expression that includes the multi-level semantic information in the text data, the visual information in the image data, the complementary association between the text data and the image data, and the potential semantic mapping relationship; using an adaptive enhancement algorithm, dynamically evaluating and adjusting the contribution of each component of the comprehensive representation vector to identify key information, wherein the key information is determined based on the multi-level semantic information in the text data, the visual information in the image data, and the complementary association between the text data and the image data, as well as the potential semantic mapping relationship; extracting content highlights in text form from the key information, and selecting representative fragments in the image data to form an information summary, wherein the content highlights are used to reflect the multi-level semantic information in the text data, and the representative fragments are used to reflect the visual information in the image data.
[0046] The technical solution of this application has the following beneficial effects:
[0047] This application can effectively obtain related text and image data from a variety of data sources, and process these data through a multimodal model, so as to achieve the complementary association between different modal information and the determination of the potential semantic mapping relationship. And by performing multi-level semantic analysis on the text and extracting visual features from the image, it is ensured that richer and deeper information can be captured. And the use of a bidirectional long short-term memory network combined with an attention mechanism to build a multimodal model helps to improve the model's ability to understand serialized information and spatially distributed information. The generator trained using a generative adversarial network is used to fuse multimodal information and generate a comprehensive representation vector containing multiple information, which provides a basis for subsequent key information identification. The adaptive enhancement algorithm is used to evaluate and adjust the contribution of different parts of the comprehensive representation vector, which improves the flexibility and accuracy in the key information identification process.
[0048] Furthermore, by initializing a set of weak classifiers and gradually updating their weights, this method can more accurately evaluate the importance of each component of the comprehensive representation vector, thereby helping to accurately locate key information. With the weight adjustment caused by error rate feedback during each round of iteration, the system can continue to learn and improve. And combining all weighted weak classifiers to form a strong classifier allows the importance of each component to be fully considered from a global perspective, which helps to obtain a more balanced and comprehensive view of key information.
[0049] Furthermore, based on the evaluation results provided by the strong classifier, important components are further screened out, and hierarchical semantic parsing technology and visual attention mechanism are applied to deeply explore the information hidden behind the text and images. The refined semantic information is integrated with the visual information through the cross-modal fusion algorithm to create a high-quality multimodal representation, which is very helpful for understanding complex scenes. Using the graph model to structure the multimodal information helps to discover elements with significant consistency between different modalities, thereby extracting the key information in the true sense and enhancing the quality and practicality of the information summary. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0051] Figure 1 A flowchart of a web page parsing method based on artificial intelligence provided in an embodiment of the present application;
[0052] Figure 2 A schematic diagram of the structure of an artificial intelligence-based web page parsing system provided in an embodiment of the present application. DETAILED DESCRIPTION
[0053] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0054] In some of the processes described in the specification and claims of this application and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent the order of precedence, and do not limit the "first" and "second" to be different types.
[0055] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0056] The specific application scenarios involved in this embodiment are all built into the web page.
[0057] Figure 1 A flowchart of a web page parsing method based on artificial intelligence is provided for an embodiment of the present application, such as Figure 1 As shown, the method includes:
[0058] 101. Obtaining associated text data and image data from multiple data sources;
[0059] This step involves collecting relevant text and image data from different sources. This data may come from social media, news articles, research web pages, etc.
[0060] In the field of news reporting, it can be articles and their accompanying pictures crawled from major news websites; in e-commerce, it can be product description text plus product pictures.
[0061] 102. Analyze the text data to capture multi-level semantic information in the text data to obtain serialization information of the text data, and extract visual features from the image data to identify visual information in the image data to obtain spatial distribution information of the image data, wherein the visual information includes key objects and key scenes;
[0062] In this step, NLP technology is used to parse the text content and identify its deep meaning, and CV technology is used to extract information about key objects or scenes from the image.
[0063] In the tourism recommendation system, sentiment analysis of user comments (text) is used to understand the quality of the travel experience, and image recognition technology is used to analyze uploaded photos and determine the characteristics of the attractions.
[0064] Optionally, the step 102 of "analyzing the text data, capturing multi-level semantic information in the text data to obtain serialization information of the text data, and performing visual feature extraction on the image data to identify visual information in the image data to obtain spatial distribution information of the image data" includes: using natural language processing technology to encode the text data to obtain a word vector representation; further processing the word vector representation through a sequence modeling method to generate serialization information of the text data; using a convolutional neural network to perform fine-grained visual feature extraction on the image data to identify visual information in the image data, wherein the visual information includes key objects and key scenes; locating and extracting the specific positions and boundaries of the key objects and key scenes to obtain spatial distribution information of the image data.
[0065] In step 102, the text data is first encoded using natural language processing (NLP) technology to generate word vector representations, and then these word vectors are further processed through sequence modeling methods to obtain the serialized information of the text; for image data, convolutional neural networks (CNNs) are used to extract fine-grained visual features, identify key objects and scenes in the image, and locate the specific positions and boundaries of these elements, thereby obtaining the spatial distribution information of the image. This method combines the advantages of text analysis and image understanding, aiming to fully capture the core semantics of multimedia content.
[0066] First, use a pre-trained language model such as BERT or Word2Vec to convert the original text into a series of numerical word vectors, each of which represents the meaning of a word. Then, sequence modeling is performed through a recurrent neural network (RNN) or its variants such as LSTM, GRU, etc., so that the model can understand the relationship between words within the sentence and the contextual meaning. At the same time, for image data, a convolutional neural network (CNN) is applied to automatically learn and extract important features in the image, such as edges, textures, etc., and gradually abstract higher-level features such as object shape, color combination, etc. through subsequent layers. Finally, specific algorithms such as target detection networks (such as YOLO, FasterR-CNN) are used to determine the exact location of key objects in the image and their bounding boxes to complete the construction of spatial distribution information.
[0067] In the embodiment of the present application, assuming that in the field of medical health, an intelligent diagnosis assistance system can be considered as an implementation case. The system first receives the medical record web page (text data) input by the doctor, converts it into a word vector through NLP technology, and then parses the logical structure and condition description of the entire document through the LSTM model. At the same time, the system also receives medical imaging data (image data) from the patient, and uses the pre-trained CNN model to identify abnormal areas (key objects) and lesion types (key scenes) from X-rays or CT scans. Through further target detection technology, the location and range of the lesion are accurately marked, and finally the text analysis results and image recognition results are integrated to provide support for clinical decision-making. This not only improves the diagnostic efficiency, but also enhances the diagnostic accuracy.
[0068] 103. Constructing a multimodal model based on a bidirectional long short-term memory network combined with an attention mechanism, wherein the multimodal model can simultaneously process the serialization information of the text data and the spatial distribution information of the image data to determine the complementary association and potential semantic mapping relationship between the text data and the image data;
[0069] In this step, a bidirectional long short-term memory network (BiLSTM) is used in combination with an attention mechanism to create a model that can integrate text serialization information and image spatial distribution information.
[0070] In the embodiment of the present application, the online education platform can use this model to evaluate students' understanding and learning progress based on the study notes (text) and classroom blackboard photos (images) submitted by the students.
[0071] Optionally, in step 103, "constructing a multimodal model based on a bidirectional long short-term memory network combined with an attention mechanism, wherein the multimodal model can simultaneously process the serialization information of the text data and the spatial distribution information of the image data to determine the complementary association and potential semantic mapping relationship between the text data and the image data" includes: using a bidirectional long short-term memory network to encode the serialization information of the text data and the spatial distribution information of the image data respectively to obtain text features and image features; dynamically adjusting the weights of the bidirectional long short-term memory network output according to the text features and image features output by the bidirectional long short-term memory network, and through a cross-attention mechanism, interacting the text features with the image features so that the text features can affect the representation of the image features, and the image features can also affect the representation of the text features, so as to establish a complementary association between the text data and the image data; using a multi-layer perceptron, integrating the serialization information of the text data and the spatial distribution information of the image data to form a cross-modal representation, wherein the cross-modal representation can reflect the potential semantic mapping relationship between the text data and the image data; and constructing a multimodal model based on the above process.
[0072] In step 103, the construction of the multimodal model aims to process the serialization information of text data and the spatial distribution information of image data by combining a bidirectional long short-term memory network (BiLSTM) and an attention mechanism. This method allows the model to simultaneously capture the complementary associations and potential semantic mapping relationships between two different types of data. Among them, BiLSTM is used to encode text sequences from a temporal dimension and can consider contextual information; and for image data, a similar strategy is used to extract its spatial features. The cross-attention mechanism allows text features and image features to influence each other, thereby enhancing the information interaction between the two, and finally a multi-layer perceptron is used to integrate these cross-modal information to form a unified representation.
[0073] First, BiLSTM is used to encode text data and image data respectively to generate their own feature vectors. Then, the attention mechanism is introduced to dynamically adjust the importance weights of the two features, especially through cross-attention to allow text features to participate in adjusting the expression of image features, and vice versa, to promote understanding and integration between different modalities. In the final stage, a multi-layer perceptron is used to combine the processed text and image features to form a new feature vector that integrates the two aspects of information and can reflect the complex semantic connection between the original inputs. The entire process ensures that the model can not only understand each type of data individually, but also identify the intrinsic connection between them.
[0074] In the embodiment of the present application, suppose that in the field of social media content analysis, it is assumed that an application that can automatically generate personalized story summaries for users is developed. The application receives travel logs (text data) and related photos (image data) from users. First, the system uses BiLSTM to encode each sentence in the travel log to capture the story line of the entire journey; at the same time, for the uploaded photos, BiLSTM (or CNN+RNN architecture adapted to images) is also applied to extract key visual elements. Then, through the cross-attention mechanism, specific words in the text description can guide image recognition to pay more attention to certain areas, such as enhancing the recognition accuracy of beach scenes when "beach" is mentioned. Conversely, if obvious landmark buildings appear in the picture, it will also prompt the text analysis part to mention more related place names. Finally, all this information is fed into a multi-layer perceptron for integration, and a complete travel story overview containing both text descriptions of wonderful moments and selected photos is output. This not only helps users quickly review their travel experiences, but also provides other readers with vivid and interesting reading materials.
[0075] 104. Based on the generator obtained by generative adversarial network training, the complementary association and potential semantic mapping relationship output by the multimodal model are fused to generate a comprehensive representation vector, wherein the comprehensive representation vector refers to a high-dimensional vector space expression including multi-level semantic information in the text data, visual information in the image data, complementary associations between the text data and the image data, and potential semantic mapping relationships;
[0076] In this step, based on the generator part in the generative adversarial network (GAN), the various information obtained in the previous step is fused into a high-dimensional vector representation.
[0077] In an embodiment of the present application, the automatic movie trailer generation system can generate a trailer containing a plot summary and visual highlights by analyzing the script text and the shot images.
[0078] Optionally, in step 104, "based on the generator trained by the generative adversarial network, the complementary associations and potential semantic mapping relationships output by the multimodal model are fused to generate a comprehensive representation vector", including: based on the generator in the generative adversarial network, the complementary associations and potential semantic mapping relationships from the multimodal model are received as input, and the input complementary associations and potential semantic mapping relationships are subjected to nonlinear transformation and deep learning processing by using the multi-layer neural network structure inside the generator, so as to achieve effective fusion of the multi-level semantic information and the visual information, while retaining the complementary associations and potential semantic mapping relationships between the text data and the image data; outputting a comprehensive representation vector through the generator, wherein the comprehensive representation vector not only includes the multi-level semantic information in the text data and the visual information in the image data, but also reflects the complementary associations and potential semantic mapping relationships between the text data and the image data.
[0079] In step 104, based on the generator part in the generative adversarial network (GAN), the model is used to receive the complementary associations and potential semantic mapping relationships between the text and the image output by the multimodal model, and perform nonlinear transformation and deep learning processing on them through a complex neural network structure. This process aims to effectively fuse the multi-level semantic information of the text data and the visual information of the image data while maintaining the established relationship between the two. Ultimately, a comprehensive representation vector is generated, which not only contains the core content of the original text and image, but also reflects the complex interaction between them.
[0080] First, the generator receives the output from the previously constructed multimodal model, which already contains text serialization information, image spatial distribution information, and the relationship between them. Then, using the complex architecture composed of multi-layer neural networks inside the generator, a series of nonlinear transformation operations are performed on these inputs to achieve effective information integration. In this step, the generator attempts to retain and enhance key features that help understand the entire multimedia content. Finally, the generator outputs a high-dimensional vector, namely the comprehensive representation vector, which can fully reflect the information of the input text and image and their relationship, providing strong support for subsequent tasks such as classification, retrieval or generation.
[0081] In the embodiment of the present application, it is assumed that in the field of news reporting automation, an automated news summary system can be considered as an implementation case. The system first extracts key information from news articles (text data) and their accompanying pictures (image data), including text content such as event descriptions, characters, and important scenes or objects appearing in the pictures. Then, the aforementioned multimodal model is used to identify the connection between these texts and images, such as the name of an important person who is mentioned in the text and appears in the picture at the same time. After that, this information is further processed by the trained GAN generator to generate a comprehensive representation vector, which not only contains the main points of the article and the significant elements in the picture, but also reflects the correlation between the two. For example, if the article mentions a political rally and the accompanying photo shows a scene of a crowd gathering, the generated vector will emphasize this point, thereby helping to automatically generate a news summary that contains both a text summary and related pictures, so that readers can quickly obtain the core content of the news.
[0082] 105. Using an adaptive enhancement algorithm, dynamically evaluate and adjust the contribution of each component of the comprehensive representation vector to identify key information, where the key information is determined based on multi-level semantic information in the text data, visual information in the image data, and complementary associations and potential semantic mapping relationships between the text data and the image data;
[0083] In this step, an adaptive boosting algorithm such as the AdaBoost method is used to automatically adjust the importance weights of each part in the comprehensive representation vector according to the current task requirements.
[0084] Using the AdaBoost method, a weak learner is trained from the training set with the initial weights. The weights of the training samples are updated according to the learning error rate of the weak learner, so that the weights of the training sample points with high learning error rates in the previous weak learner become higher. Then these points with high error rates are given higher attention in the weak learner, and the weak learner is trained using the training set with adjusted weights. This is repeated until the number of weak learners reaches the pre-specified number T, and finally these T weak learners are integrated through the set strategy to obtain the final strong learner.
[0085] In the embodiment of the present application, the personalized advertising push service will continuously optimize the recommendation algorithm based on the user's browsing history (text) and click behavior (image) to improve the relevance and attractiveness of the advertisement.
[0086] Optionally, the step 105 of "using the adaptive enhancement algorithm to dynamically evaluate and adjust the contribution of each component of the comprehensive representation vector to identify key information" includes: initializing a group of weak classifiers using the adaptive enhancement algorithm, performing preliminary evaluation processing on each component of the comprehensive representation vector, and obtaining the initial weight of each weak classifier; updating the weight of each weak classifier according to the error rate of the current round to obtain an updated weight; retraining each weak classifier based on the updated weight, and calculating a new weight coefficient to obtain a new weight coefficient reflecting the importance of each weak classifier in the final decision; combining all weighted weak classifiers to form a strong classifier, performing a global evaluation processing on each component of the comprehensive representation vector, and obtaining an evaluation result of the strong classifier; determining the contribution of each component of the comprehensive representation vector according to the evaluation result of the strong classifier; performing a comprehensive analysis processing according to the contribution of each component of the comprehensive representation vector, combined with the multi-level semantic information in the text data, the visual information in the image data, and the complementary association between the text data and the image data and the potential semantic mapping relationship, to determine the key information.
[0087] In step 105, an adaptive boosting algorithm (such as AdaBoost) is used to dynamically evaluate and adjust the importance of each component of the comprehensive representation vector to identify key information. This method processes different parts of the comprehensive representation vector by initializing a set of weak classifiers, and continuously updates the weights of each weak classifier based on its performance, and finally combines these weighted weak classifiers to form a strong classifier. This process enables the system to automatically identify the features that are most important for decision-making, thereby determining which are the key information in the comprehensive representation vector.
[0088] First, multiple weak classifiers are initialized using an adaptive boosting algorithm and initial weights are assigned to them. Then, the weight of each weak classifier is updated based on the error rate of the current round; then, each weak classifier is retrained according to the updated weights and a new weight coefficient is calculated, which reflects the importance of each weak classifier in the final decision. Next, all weighted weak classifiers are combined to form a strong classifier for a comprehensive evaluation of all components of the comprehensive representation vector. Finally, the contribution of each part of the comprehensive representation vector is determined based on the results given by the strong classifier, and the key information is identified by combining the complementarity and semantic mapping relationship between text and image data.
[0089] Among them, in this embodiment, the AdaBoost algorithm only trains the same basic classifier (weak classifier) for different training sets, and then combines these classifiers obtained on different training sets to form a stronger final classifier (strong classifier). Theoretically, as long as the classification ability of each weak classifier is better than random guessing, when its number tends to infinity, the error rate of the strong classifier will tend to zero. Different training sets in the AdaBoost algorithm are only achieved by adjusting the weight corresponding to each sample. At the beginning, the weight corresponding to each sample is the same, and a basic classifier h, (x) is trained under this sample distribution. For samples that are misclassified by h, (x), the weight of the corresponding sample is increased; and for samples that are correctly classified, its weight is reduced. In this way, the misclassified samples can be highlighted and a new sample distribution can be obtained. At the same time, h, (x) is given a weight according to the misclassification situation, indicating the importance of the basic classifier. The less misclassification, the greater the weight. Under the new sample distribution, the basic classifier is trained again to obtain the basic classifier h, (x) and its weight. By analogy, after T such cycles, we get T basic classifiers and T corresponding weights. Finally, add up these T basic classifiers according to certain weights to get the desired strong classifier.
[0090] In an embodiment of the present application, it is assumed that in an intelligent customer service system, a scenario can be considered in which the system needs to quickly locate the core of the problem from the problem description (text data) submitted by the user and the related screenshots or photos (image data). First, a comprehensive representation vector is obtained by applying a multimodal model to the information provided by the user, which contains multi-level semantic information described in the text and visual information in the image. Then, a set of weak classifiers are initialized using the AdaBoost algorithm to perform preliminary evaluations on different aspects of the comprehensive representation vector. With each round of iteration, the weights of the weak classifiers are dynamically adjusted according to their performance until a strong classifier is constructed. This strong classifier can accurately point out which words or image elements are most critical to understanding user problems. For example, if a user mentions "account login failed" and attaches a screenshot of the error page, the system will identify keywords such as "login", "failure" and the specific image area of the error code as key information for solving the problem. In this way, customer service personnel or automated processes can directly take action on these key points to improve service efficiency and accuracy.
[0091] Optionally, the step 105 of "comprehensively analyzing and processing to determine key information based on the contribution of each component of the comprehensive representation vector, combined with the multi-level semantic information in the text data, the visual information in the image data, and the complementary association between the text data and the image data and the potential semantic mapping relationship" includes: using the evaluation result of the strong classifier to quantify the contribution of each component of the comprehensive representation vector to obtain an importance score for each component; based on the importance score, screening out components whose contribution meets the set conditions, and for the screened components, combining hierarchical semantic parsing technology to deeply explore the deep meaning in the text data, and using the visual attention mechanism to further identify the visual information in the image data to obtain refined multi-level semantic information and visual information; integrating the refined semantic information with the visual information through a cross-modal fusion algorithm to generate a multimodal representation; based on the multimodal representation, using a graph model to structure the refined semantic information and visual information, identify elements with significant consistency between different modalities, and obtain key information.
[0092] In step 105, the components in the comprehensive representation vector are quantified by the evaluation results of the strong classifier to determine their importance scores. Then, based on these scores, the parts with higher importance are screened out, and hierarchical semantic parsing technology and visual attention mechanism are further used to deeply explore the deep meaning of text and image data. Subsequently, the cross-modal fusion algorithm is used to integrate the refined information into a unified multimodal representation. Finally, a graphical model is used to structure this information and identify elements with significant consistency between different modalities, thereby determining key information. This approach ensures that the most relevant and representative content is extracted from a large amount of data.
[0093] First, score each component in the comprehensive representation vector based on the output of the strong classifier to reflect its importance to the overall decision. Then, according to the set criteria, select the parts with higher scores as the focus of subsequent analysis. Next, use hierarchical semantic parsing technology to deeply understand the complex semantic structure in the text data, and combine the visual attention mechanism to strengthen the recognition of key visual features in the image data. On this basis, the two aspects of information are combined through the cross-modal fusion algorithm to form a more comprehensive multimodal representation. The last step is to build a graph model. By modeling the relationship between the elements in the multimodal representation, find the core information points that show a high degree of consistency between the text and the image, so as to determine the final key information.
[0094] In the embodiment of the present application, it is assumed that in the intelligent legal document review system, such a scenario can be considered: the system needs to quickly locate important terms and diagram contents from the contract text (text data) and its accompanying schematic diagram or flowchart (image data). First, the system uses the comprehensive representation vector obtained in the previous step and evaluates it through a strong classifier to score each component. Assuming that the scores of "payment terms" and "delivery flowchart" are high, they are selected as the focus of subsequent analysis. Then, hierarchical semantic parsing technology is applied to deeply understand the specific details of "payment terms", such as payment methods, time arrangements, etc.; at the same time, the visual attention mechanism is used to focus on the key nodes and paths involved in the "delivery flowchart". After that, these refined information are integrated together through a cross-modal fusion algorithm to form a multimodal representation containing comprehensive information of text and image. Finally, a graph model is established to explore the correspondence between text description and image display. If it is found that the "payment time node" is closely related to the "specific stage on the delivery flowchart", then this part of information is identified as key information in the contract review. In this way, auditors can quickly grasp the core points of the contract and improve work efficiency.
[0095] It should be noted that "selecting components whose contributions meet the set conditions based on importance scores, and combining hierarchical semantic parsing technology and visual attention mechanism to further process text and image data" has several key purposes:
[0096] By filtering out components with high contribution, computing resources can be focused on the information that is most valuable for decision making. This not only improves processing speed, but also reduces interference from irrelevant or low-value information.
[0097] For the important components selected, more in-depth techniques (such as hierarchical semantic parsing and visual attention mechanism) are used for analysis to reveal the deeper meaning and details behind these parts. For example, in text, hierarchical semantic parsing can help understand complex language features such as sentence structure and contextual relationships; in images, visual attention mechanisms help identify which areas or objects are key elements in the image.
[0098] When key information extracted from different modalities (such as text and images) is integrated, it becomes particularly important to ensure that there is significant consistency between these information. By first refining the information within each modality, the connection points between them can be better captured, thus promoting the effectiveness of cross-modal fusion.
[0099] By focusing on highly contributing information and processing it in a refined manner, a more accurate and comprehensive multimodal representation can be generated, which is crucial for subsequent steps (such as using graph models for structured processing). Doing so helps the system identify key information that is highly consistent across multiple data types, thereby making more reliable judgments or predictions.
[0100] It should also be noted that in the process of "integrating the refined semantic information with the visual information through a cross-modal fusion algorithm to generate a multimodal representation; based on the multimodal representation, using a graph model to structure the refined semantic information and visual information, identifying elements with significant consistency between different modalities, and obtaining key information", although it seems that there is a component of repeated processing, in fact, each stage is contributing to the ultimate goal of identifying key information. After the refined semantic information and visual information are cross-modally fused to generate a multimodal representation, they are then structured through a graph model in order to further reveal the deep relationship between the information, thereby improving the system's understanding and decision-making capabilities. This process is not a simple repetition, but a process of gradual deepening and refinement.
[0101] This application takes into account that in machine learning, especially for classification problems, AdaBoost (AdaptiveBoosting) is a very effective ensemble learning method. It iteratively trains weak classifiers and dynamically adjusts sample weights according to the performance of each classifier, and finally combines these weak classifiers into a strong classifier. The above formula describes the key steps of the AdaBoost algorithm, including the process of initializing sample weights, training weak classifiers, updating weak classifier weights, updating sample weights, and combining weak classifiers to form a strong classifier. In addition, a method for calculating the importance scores of each component of the comprehensive representation vector is introduced.
[0102] The options are as follows:
[0103] Optionally, the step 105 of “using the evaluation result of the strong classifier to quantify the contribution of each component of the comprehensive representation vector to obtain an importance score of each component” includes:
[0104] Initialize sample weights:
[0105] ;
[0106] in, is the initial weight of the i-th sample, and N is the total number of samples;
[0107] Train a weak classifier and calculate the error rate:
[0108] ;
[0109] in, is the error rate of the mth weak classifier, represents the sum operation from 1 to N,
[0110] is the weight of the i-th sample after the m-1th iteration, N is the total number of samples, It is an indicator function. It takes the value 1 when the condition in the brackets is true, otherwise it takes the value 0. is the true label of the i-th sample, is the predicted label of the mth weak classifier for the i-th sample;
[0111] Update the weak classifier weights:
[0112] ;
[0113] in, is the importance score of the mth weak classifier, is the error rate of the mth weak classifier, is a regularization parameter used to prevent overfitting. The logarithm function is used to calculate the natural logarithm of two ratios;
[0114] Update sample weights:
[0115] ;
[0116] in, is the weight of the i-th sample after the m-th iteration, is the weight of the i-th sample after the m-1th iteration, is an exponential function term, used to adjust the weight of the sample. is the importance score of the mth weak classifier, is the true label of the i-th sample, is the predicted label of the mth weak classifier for the i-th sample, N is the total number of samples, is a normalization factor that ensures that the sum of all sample weights is 1:
[0117] ;
[0118] Combine weak classifiers to form a strong classifier:
[0119] ;
[0120] Among them, H(x) is the final strong classifier, is the importance score of the mth weak classifier, is the predicted label of the mth weak classifier for the i-th sample, represents the sum operation from 1 to M, where M is the number of weak classifiers, is a bias term used to adjust the classification threshold;
[0121] Compute the importance scores of the components of the composite representation vector:
[0122] ;
[0123] in, is the importance score of the jth component in the comprehensive representation vector, is the contribution of the mth weak classifier to the jth component, is the regularization coefficient, is the regularization term of the jth component, which is used to control the model complexity.
[0124] Sample weight initialization: ensure that each sample is treated equally at the beginning.
[0125] Error rate calculation: Evaluate the performance of each weak classifier so that its weight can be adjusted later.
[0126] Weak classifier weight update: Determine the influence of each weak classifier on the final decision based on the error rate.
[0127] Sample weight update: Redistribute sample weights based on the performance of the current weak classifier so that difficult-to-classify samples receive higher weights.
[0128] Strong classifier combination: Use weighted voting to integrate the results of all weak classifiers to form a more powerful classification capability.
[0129] Importance score calculation: measures the importance of each component in the comprehensive representation vector to the overall decision.
[0130] The following is a brief introduction to the design reasons of each sub-item of the formula:
[0131] ;
[0132] Sample weight initialization At the beginning of the AdaBoost algorithm, all samples are considered equally important. , where N is the total number of samples), ensuring that the model is not biased towards any particular sample in the first round of iterations.
[0133] The following is a brief introduction to how to obtain the parameters of the formula:
[0134] N (the total number of samples) is obtained directly from the dataset. For example, if the dataset contains 1000 samples, then N=1000.
[0135] The following is a brief introduction to the design of this formula:
[0136] ;
[0137] The error rate reflects the performance of the weak classifier. By using a weighted error rate, the importance of samples that were assigned higher weights in the previous round can be emphasized. This encourages the weak classifiers trained subsequently to pay more attention to these difficult-to-classify samples.
[0138] The following is a brief introduction to how to obtain the parameters of the formula:
[0139] in, (true label), (Feature vector) is obtained directly from the training dataset. Each sample has a corresponding true label and feature vector. (Weak classifier prediction result) is predicted by the trained weak classifier model. For example, decision trees, logistic regression, etc. can be used as weak classifiers.
[0140] The following is a brief introduction to the design of this formula:
[0141] ;
[0142] This formula adjusts the weights of weak classifiers based on the error rate. When the error rate is low, The larger the value, the more important the weak classifier is in the final decision; vice versa. This prevents numerical instabilities in extreme cases and helps avoid overfitting.
[0143] The following is a brief introduction to how to obtain the parameters of this formula:
[0144] in, (Regularization parameter) is used to prevent overfitting, and the best value is usually determined by cross-validation. For example, a suitable value can be selected between 0.01 and 1.
[0145] The following is a brief introduction to the design of this formula:
[0146] ;
[0147] This process dynamically adjusts the sample weights based on the performance of the current weak classifier. If a sample is correctly classified, its weight decreases; if it is misclassified, its weight increases. In this way, subsequent weak classifiers will pay more attention to those samples that were previously misclassified. Normalization factor Make sure the sum of the updated sample weights is still 1.
[0148] The following is a brief introduction to the design of this formula:
[0149] ;
[0150] A strong classifier is formed by weighted summing up the prediction results of all weak classifiers. The importance score of each weak classifier Determines its influence in the final decision. It can be used to adjust the classification threshold to optimize model performance.
[0151] The following is a brief introduction to how to obtain the parameters of this formula:
[0152] Among them, M (number of weak classifiers) is a hyperparameter, and the best value is usually selected through cross-validation. It can also be pre-set based on experience or domain knowledge.
[0153] The following is a brief introduction to the design of this formula:
[0154] ;
[0155] This product reflects the combination of the importance score of the mth weak classifier on the jth feature or component and its contribution. In this way, it is possible to identify which features or components play a key role in the entire model. The purpose of introducing regularization terms is to control model complexity and prevent overfitting. and the regularization term , which can balance the generalization ability and fitting accuracy of the model.
[0156] The following is a brief introduction to how to obtain the parameters of the formula:
[0157] in, (Regularization coefficient) is also a hyperparameter used to control model complexity and prevent overfitting. It can be optimized through cross-validation. (Regularization term) Usually L1 or L2 regularization term, depending on the application scenario. L1 regularization helps sparsity, L2 regularization helps smoothness.
[0158] In the embodiment of the present application, it is assumed that there is a data set for a binary classification problem, which contains 4 samples, each of which has 3 features. The AdaBoost algorithm is used to construct a strong classifier and calculate the importance score of each component of the comprehensive representation vector.
[0159] Dataset Sample ,sample ,sample ,sample .
[0160] Initialize sample weights:
[0161] ;
[0162] Assume that a simple threshold classifier is used as a weak classifier. For feature 1, the threshold is 5:
[0163] ;
[0164] Error rate: ;
[0165] Update the weak classifier weights:
[0166] ;
[0167] Update sample weights:
[0168] Normalization factor: ;
[0169] New sample weights:
[0170] ;
[0171] Repeat the above steps until all weak classifiers are trained.
[0172] Suppose we train a weak classifier based on feature 2 with a threshold of 8:
[0173] ;
[0174] Error rate: ;
[0175] Weak classifier weights: ;
[0176] Update sample weights (omit the specific calculation process)...
[0177] Strong classifier: ;
[0178] Assume the following contribution:
[0179] ;
[0180] Regularization parameter: ;
[0181] Regularization term: Assumption ;
[0182] Importance Rating:
[0183] ;
[0184] Through the above calculations, we can see the importance scores of different features in the final model. In this example, the importance score of feature 1 is 0.07 (a positive number), indicating that it has a certain importance in the current model; while the importance scores of feature 2 and feature 3 are -0.07 and -0.04 (negative numbers), respectively, indicating that they play a small role in the current model. This analysis method not only helps to understand which features are most critical for classification tasks, but also provides a basis for subsequent feature selection and model optimization.
[0185] 106. Extracting key points of content in text form from the key information, and selecting representative fragments in the image data to form an information summary, wherein the key points of content are used to reflect the multi-level semantic information in the text data, and the representative fragments are used to reflect the visual information in the image data.
[0186] In this step, the most core information points are finally extracted from the processed data and presented in a concise and clear form.
[0187] In the embodiment of the present application, for a research paper review writing tool, the main points of the content such as the research background and methodological innovations can be automatically summarized from a large amount of literature, and representative experimental result figures can be selected as supplementary explanations.
[0188] Optionally, the step 106 of "extracting key points of content in text form from the key information, and selecting representative fragments from the image data to form an information summary" includes: extracting key points of content in text form from the key information using a text summary algorithm; selecting representative fragments from the image data based on image segmentation technology; generating an information summary based on the key points of content and the representative fragments, wherein the key points of content are used to reflect the multi-level semantic information in the text data, and the representative fragments are used to reflect the visual information in the image data.
[0189] In step 106, the key points of the content in the form of text are extracted from the key information, and representative fragments in the image data are selected to form an information summary. This process includes using a text summary algorithm to extract the core content of the text, and using image segmentation technology to select the fragments that best represent the main information of the image. In this way, the generated information summary can not only accurately reflect the multi-level semantic information in the text data, but also supplement the description through visual elements, making the summary more comprehensive and intuitive.
[0190] First, a text summary algorithm (such as an extractive or generative summary method) is used to process the previously determined key information to extract text points that can summarize the main content. Then, image segmentation technology (such as a deep learning-based method) is used to analyze the image data to identify and select the most representative image areas or objects as part of the image summary. Finally, these text points are combined with the selected image fragments to form an information summary that contains both text descriptions and visual presentations, ensuring that users can quickly understand the main content of the original material and its visual representation.
[0191] In an embodiment of the present application, it is assumed that in a news reporting automation system, the system needs to automatically generate a concise and attractive summary for each report. Suppose a report tells about a major natural disaster that occurred in a certain place, which contains a detailed description of the event, rescue operations, and photos of the disaster situation. First, the system uses a text summary algorithm to extract core information such as the scope of the disaster, the number of casualties, and the progress of the rescue from the report; then, through image segmentation technology, key pictures such as the severity of the disaster and the scenes of the rescue team's work are selected from the accompanying photos. Finally, combining these text key points and selected pictures, an information summary is generated, so that readers can not only understand the full picture of the incident through text, but also intuitively feel the situation on the scene through pictures, so as to more comprehensively grasp the key content of the report.
[0192] Figure 2 A schematic diagram of the structure of a web page parsing system based on artificial intelligence is provided for the embodiment of the present application, such as Figure 2 As shown, the device comprises:
[0193] The acquisition module 21 is used to acquire associated text data and image data from multiple data sources; analyze the text data to capture multi-level semantic information in the text data to obtain serialization information of the text data, and extract visual features from the image data to identify visual information in the image data to obtain spatial distribution information of the image data;
[0194] A construction module 22 is used to construct a multimodal model based on a bidirectional long short-term memory network combined with an attention mechanism, wherein the multimodal model can simultaneously process the serialization information of the text data and the spatial distribution information of the image data to determine the complementary association and potential semantic mapping relationship between the text data and the image data;
[0195] A fusion module 23 is used to fuse the complementary associations and potential semantic mapping relationships output by the multimodal model based on a generator trained by a generative adversarial network to generate a comprehensive representation vector, wherein the comprehensive representation vector refers to a high-dimensional vector space expression including multi-level semantic information in the text data, visual information in the image data, complementary associations between the text data and the image data, and potential semantic mapping relationships;
[0196] A recognition module 24, configured to dynamically evaluate and adjust the contribution of each component of the comprehensive representation vector using an adaptive enhancement algorithm to identify key information, wherein the key information is determined based on multi-level semantic information in the text data, visual information in the image data, and complementary associations and potential semantic mapping relationships between the text data and the image data;
[0197] The generation module 25 is used to extract the main points of the content in text form from the key information, and select representative fragments in the image data to form an information summary. The main points of the content are used to reflect the multi-level semantic information in the text data, and the representative fragments are used to reflect the visual information in the image data.
[0198] Figure 2 The web page parsing system based on artificial intelligence can perform Figure 1 The implementation principle and technical effect of the web page parsing method based on artificial intelligence described in the embodiment shown are not repeated here. The specific way in which each module and unit performs operations in the web page parsing system based on artificial intelligence in the above embodiment has been described in detail in the embodiment of the method, and will not be elaborated here.
[0199] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0200] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0201] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A web page parsing method based on artificial intelligence, characterized in that: include: Acquire associated text data and image data from various data sources; Analyzing the text data to capture multi-level semantic information in the text data to obtain serialization information of the text data, and extracting visual features from the image data to identify visual information in the image data to obtain spatial distribution information of the image data, wherein the visual information includes key objects and key scenes; A multimodal model is constructed based on a bidirectional long short-term memory network combined with an attention mechanism, wherein the multimodal model can simultaneously process the serialization information of the text data and the spatial distribution information of the image data to determine the complementary association and potential semantic mapping relationship between the text data and the image data; Based on the generator obtained by generative adversarial network training, the complementary association and potential semantic mapping relationship output by the multimodal model are fused to generate a comprehensive representation vector, wherein the comprehensive representation vector refers to a high-dimensional vector space expression including multi-level semantic information in the text data, visual information in the image data, complementary associations between the text data and the image data, and potential semantic mapping relationships; Using an adaptive enhancement algorithm, dynamically evaluating and adjusting the contribution of each component of the comprehensive representation vector to identify key information, wherein the key information is determined based on multi-level semantic information in the text data, visual information in the image data, and complementary associations and potential semantic mapping relationships between the text data and the image data; The key information is extracted from the key information in text form, and representative fragments in the image data are selected to form an information summary. The key points are used to reflect the multi-level semantic information in the text data, and the representative fragments are used to reflect the visual information in the image data.
2. The method according to claim 1, characterized in that The adaptive enhancement algorithm is used to dynamically evaluate and adjust the contribution of each component of the comprehensive representation vector to identify key information, including: Initializing a group of weak classifiers using an adaptive boosting algorithm, performing preliminary evaluation processing on each component of the comprehensive representation vector, and obtaining an initial weight of each weak classifier; According to the error rate of the current round, the weights of each weak classifier are updated to obtain updated weights; Based on the updated weights, each weak classifier is retrained and a new weight coefficient is calculated to obtain a new weight coefficient that reflects the importance of each weak classifier in the final decision; Combining all weighted weak classifiers to form a strong classifier, performing global evaluation processing on each component of the comprehensive representation vector, and obtaining an evaluation result of the strong classifier; Determining the contribution of each component of the comprehensive representation vector according to the evaluation result of the strong classifier; According to the contribution of each component of the comprehensive representation vector, combined with the multi-level semantic information in the text data, the visual information in the image data, the complementary association between the text data and the image data, and the potential semantic mapping relationship, comprehensive analysis and processing are performed to determine the key information.
3. The method according to claim 2, characterized in that According to the contribution of each component of the comprehensive representation vector, combined with the multi-level semantic information in the text data, the visual information in the image data, and the complementary association and potential semantic mapping relationship between the text data and the image data, a comprehensive analysis process is performed to determine key information, including: Quantifying the contribution of each component of the comprehensive representation vector using the evaluation result of the strong classifier to obtain an importance score for each component; Based on the importance score, the components whose contribution meets the set conditions are screened out, and for the screened out components, the hierarchical semantic parsing technology is combined to deeply explore the deep meaning in the text data, and the visual attention mechanism is used to further identify the visual information in the image data, so as to obtain refined multi-level semantic information and visual information; Through the cross-modal fusion algorithm, the refined semantic information is integrated with the visual information to generate a multimodal representation; Based on the multimodal representation, a graphical model is used to perform structured processing on the refined semantic information and visual information, identify elements with significant consistency between different modalities, and obtain key information.
4. The method according to claim 1, characterized in that: Analyzing the text data to capture multi-level semantic information in the text data to obtain serialization information of the text data, and extracting visual features from the image data to identify visual information in the image data to obtain spatial distribution information of the image data, including: Using natural language processing technology, the text data is encoded to obtain a word vector representation; Further processing the word vector representation by a sequence modeling method to generate serialization information of the text data; Using a convolutional neural network to perform fine-grained visual feature extraction processing on the image data to identify visual information in the image data, wherein the visual information includes key objects and key scenes; The specific positions and boundaries of the key objects and key scenes are located and extracted to obtain the spatial distribution information of the image data.
5. The method according to claim 1, characterized in that A multimodal model is constructed based on a bidirectional long short-term memory network combined with an attention mechanism. The multimodal model can simultaneously process the serialization information of the text data and the spatial distribution information of the image data to determine the complementary association and potential semantic mapping relationship between the text data and the image data, including: Using a bidirectional long short-term memory network, the serialization information of the text data and the spatial distribution information of the image data are respectively encoded to obtain text features and image features; According to the text features and image features output by the bidirectional long short-term memory network, the weight of the bidirectional long short-term memory network output is dynamically adjusted, and the text features and the image features are interacted with each other through a cross attention mechanism, so that the text features can affect the representation of the image features, and the image features can also affect the representation of the text features, so as to establish a complementary association between the text data and the image data; A multi-layer perceptron is used to integrate the serialization information of the text data and the spatial distribution information of the image data to form a cross-modal representation to complete the construction of a multimodal model, wherein the cross-modal representation can reflect the potential semantic mapping relationship between the text data and the image data.
6. The method according to claim 1, characterized in that Based on the generator obtained by generative adversarial network training, the complementary association and the potential semantic mapping relationship output by the multimodal model are fused to generate a comprehensive representation vector, including: Based on the generator in the generative adversarial network, the complementary associations and potential semantic mapping relationships from the multimodal model are received as input, and the multi-layer neural network structure inside the generator is used to perform nonlinear transformation and deep learning processing on the input complementary associations and potential semantic mapping relationships, so as to achieve effective fusion of the multi-level semantic information and the visual information, and at the same time retain the complementary associations and potential semantic mapping relationships between the text data and the image data; The generator outputs a comprehensive representation vector, which not only includes the multi-level semantic information in the text data and the visual information in the image data, but also reflects the complementary relationship and potential semantic mapping relationship between the text data and the image data.
7. The method according to claim 1, characterized in that Extracting the key information in text form, and selecting representative segments from the image data to form an information summary, including: Use text summarization algorithms to extract the main points of content in text form from key information; Selecting representative segments from the image data according to image segmentation technology; An information summary is generated according to the content highlights and the representative segments, wherein the content highlights are used to reflect the multi-level semantic information in the text data, and the representative segments are used to reflect the visual information in the image data.
8. A web page parsing system based on artificial intelligence, characterized in that: include: An acquisition module, used for acquiring associated text data and image data from various data sources; Analyzing the text data to capture multi-level semantic information in the text data to obtain serialization information of the text data, and extracting visual features from the image data to identify visual information in the image data to obtain spatial distribution information of the image data; A construction module is used to construct a multimodal model based on a bidirectional long short-term memory network combined with an attention mechanism, wherein the multimodal model can simultaneously process the serialization information of the text data and the spatial distribution information of the image data to determine the complementary association and potential semantic mapping relationship between the text data and the image data; A fusion module, for fusing the complementary associations and potential semantic mapping relationships output by the multimodal model based on a generator trained by a generative adversarial network, and generating a comprehensive representation vector, wherein the comprehensive representation vector refers to a high-dimensional vector space expression including multi-level semantic information in the text data, visual information in the image data, complementary associations between the text data and the image data, and potential semantic mapping relationships; A recognition module, for dynamically evaluating and adjusting the contribution of each component of the comprehensive representation vector using an adaptive enhancement algorithm to identify key information, wherein the key information is determined based on multi-level semantic information in the text data, visual information in the image data, and complementary associations and potential semantic mapping relationships between the text data and the image data; A generation module is used to extract content highlights in text form from the key information and select representative fragments in the image data to form an information summary. The content highlights are used to reflect the multi-level semantic information in the text data, and the representative fragments are used to reflect the visual information in the image data.
Citation Information
Patent Citations
Graph contrast learning method and device based on adaptive data enhancement and storage medium
CN116543288A
Public opinion sentiment classification method and system based on machine learning
CN118070103A