Abstract Generation Method and System for Semantic Mining of Newspaper Articles
By performing semantic extraction and matching analysis of newspaper articles, optimizing semantic space and using feature identification technology, the problem of difficulty in understanding complex semantic structures and analyzing semantic relationships between articles in the existing technology is solved, and efficient and accurate semantic correlation mining and abstract generation are achieved.
Patent Information
- Application Number
- CN202510280238.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-11
AI Technical Summary
It is difficult for the prior art to accurately understand the core content and semantic relationships of newspapers and magazines with complex semantic structures, and the lack of in-depth analysis of semantic relationships between different articles, resulting in the abstract results that cannot effectively reflect the semantic connections between multiple articles.
By semantic extraction of newspaper articles, the article semantic space is generated, and semantic matching analysis is carried out, the semantic space is optimized, the weight coefficient is generated, and the semantic correlation vector units are used to locate semantic related vector units. Finally, accurate summary results are generated through in-depth semantic correlation analysis.
It significantly improves the accuracy and efficiency of semantic correlation mining between newspapers and magazines, and the generated abstract results are more accurate and comprehensive, which promotes the application effectiveness of information retrieval, content recommendation and public opinion analysis.
Smart Images

Figure CN119807409B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology. Specifically, it relates to a method and system for generating abstracts applied to semantic mining of newspaper articles. Background Art
[0002] In the era of information explosion, newspapers, as traditional information dissemination media, still carry a large amount of valuable information. However, with the continuous increase in the number of newspaper articles, how to quickly and accurately extract key information from a vast number of articles has become an urgent problem to be solved. Especially when semantic correlation analysis needs to be performed on multiple articles, traditional methods are often inefficient and difficult to accurately capture the semantic connections between articles.
[0003] Most of the existing methods for generating abstracts of newspaper articles are based on keyword extraction or text statistical features. These methods may show certain effects when dealing with simple texts, but when facing newspaper articles with complex semantic structures, it is often difficult to accurately understand the core content and semantic associations of the articles. In addition, these methods usually lack in-depth analysis of the semantic relationships between different articles, resulting in the generated abstract results often only focusing on the content of a single article and unable to effectively reflect the semantic connections between multiple articles. Summary of the Invention
[0004] In view of the above-mentioned problems, in combination with the first aspect of this application, embodiments of this application provide a method for generating abstracts applied to semantic mining of newspaper articles, and the method includes:
[0005] Performing semantic extraction on the first newspaper article to generate a first article semantic space, and performing semantic extraction on the second newspaper article to generate a second article semantic space;
[0006] Performing semantic matching analysis on the first article semantic space to generate weight coefficients of each first semantic vector unit in the first article semantic space, and performing semantic matching analysis on the second article semantic space to generate weight coefficients of each second semantic vector unit in the second article semantic space, where the weight coefficients reflect the influence coefficients on the semantic association of newspaper articles;
[0007] Optimizing the first article semantic space according to the weight coefficients of each first semantic vector unit to generate x first semantic vector units, and optimizing the second article semantic space according to the weight coefficients of each second semantic vector unit to generate x second semantic vector units;
[0008] Perform semantic matching analysis on the x first semantic vector units and the x second semantic vector units to generate a first feature identifier corresponding to the x first semantic vector units and a second feature identifier corresponding to the x second semantic vector units. The first feature identifier is used to locate y first semantic vector units among the x first semantic vector units that are semantically relevant to the x second semantic vector units, and the second feature identifier is used to locate z second semantic vector units among the x second semantic vector units that are semantically relevant to the x first semantic vector units. Herein, the first feature identifier and the second feature identifier are used to represent corresponding masks;
[0009] Perform semantic relevance analysis on the y first semantic vector units and the z second semantic vector units to generate a first semantic association summary result between the first newspaper article and the second newspaper article.
[0010] In another aspect, an embodiment of the present application further provides a summary generation system, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor. The machine-readable storage medium is used to store programs, instructions, or codes, and the processor is used to execute the programs, instructions, or codes in the machine-readable storage medium to implement the above method.
[0011] Based on the above aspects, the embodiment of the present application significantly improves the accuracy and efficiency of semantic association mining between newspaper articles through semantic extraction and matching analysis. Specifically, first, semantic extraction is performed on two articles respectively to construct their respective semantic spaces, and through weight coefficient analysis, semantic vector units that have a significant impact on the semantic association of the articles are effectively identified. Subsequently, through semantic space optimization, the most representative semantic vector units are further refined, and the feature identifier technology is used to accurately locate the vector units with semantic relevance between the two articles. Finally, through in-depth semantic relevance analysis, a summary result that accurately reflects the semantic association relationship between the two articles is generated, which not only improves the accuracy and comprehensiveness of summary generation, but also greatly promotes the application efficiency in fields such as information retrieval, content recommendation, and public opinion analysis. Description of the Drawings
[0012] Figure 1 is a schematic execution flowchart of a summary generation method for newspaper article semantic mining provided by an embodiment of the present application.
[0013] Figure 2 is a schematic hardware architecture diagram of a summary generation system provided by an embodiment of the present application. Detailed Embodiments
[0014] The following specifically describes the present application in conjunction with the drawings of the specification, Figure 1It is a schematic flowchart of a summary generation method for newspaper article semantic mining provided by an embodiment of the present application. The summary generation method for newspaper article semantic mining will be introduced in detail below.
[0015] Step S110: Extract the semantics of the first newspaper article to generate a first article semantic space, and extract the semantics of the second newspaper article to generate a second article semantic space.
[0016] Specifically, the first newspaper article / second newspaper article is the text object to be analyzed. For example, one of them is a movie review article about a certain popular movie (the first newspaper article), and the other is an entertainment news article about the lead actor of the movie (the second newspaper article). The article semantic space is to map the semantic information of the newspaper article into a multi-dimensional vector space through semantic extraction, where each word or phrase corresponds to a vector in the space.
[0017] That is to say, in this embodiment, for the first newspaper article (movie review article), the server first analyzes each word in the first newspaper article. For example, the first newspaper article mentions the plot of the movie, such as "the protagonist encountered a mysterious magical creature during the adventure". The server will identify the semantics of words such as "protagonist", "adventure", "mysterious", and "magical creature", and map these words into a multi-dimensional semantic space, where each word corresponds to a vector in this space. For words describing the movie scene, such as "gorgeous visual effects", semantic mapping is also performed in the same way. In this way, the semantics of the entire movie review article are represented in the form of vectors in this multi-dimensional space, forming the first article semantic space. This semantic space contains various semantic information of the movie review article, including evaluations of the movie plot, scene, performance, etc.
[0018] For the second newspaper article (entertainment news article), if it mentions that "the lead actor recently participated in a charity event and shared interesting stories during the filming of the movie", the server will perform semantic extraction on words such as "lead actor", "charity event", "filming of the movie", and "interesting stories". Represent these words and their semantic relationships in the form of vectors in another multi-dimensional space to generate the second article semantic space. This second article semantic space reflects various semantic information in the entertainment news about the lead actor's activities and movie-related events, etc.
[0019] Step S120: Perform semantic matching analysis on the first article semantic space to generate weight coefficients for each first semantic vector unit in the first article semantic space, and perform semantic matching analysis on the second article semantic space to generate weight coefficients for each second semantic vector unit in the second article semantic space. The weight coefficients reflect the influence coefficients on the semantic association of the newspaper article.
[0020] In this embodiment, after the semantic vector unit is optimized in the semantic space, x most representative and important semantic vector units are selected from the article semantic space as the objects for subsequent feature identification and correlation analysis.
[0021] Taking the previous movie review article (the first newspaper article) as an example, in the first article semantic space, semantic vector units related to "movie plot" may be more important when analyzing the overall semantic correlation of the movie review. If most of the content in the movie review focuses on the plot, then the semantic vector unit of "movie plot" may be assigned a higher weight coefficient. For example, for the semantic vector unit related to the plot "The protagonist encounters a mysterious magical creature during the adventure", the server analyzes its correlation degree with other semantic units in the movie review (such as the evaluation of the protagonist's acting skills, the theme expression of the movie, etc.). If it is found that many evaluations of the acting skills are based on this plot, then this plot-related semantic vector unit has a greater impact on the semantic correlation of the entire movie review article, and the weight coefficient may be relatively high.
[0022] For entertainment news articles (the second newspaper article), if the semantic vector unit of "the lead actor shares interesting stories during the filming of the movie" has a strong correlation with other information about the lead actor in the news (such as the image shaping of the lead actor, public image, etc.), then its weight coefficient in the second article semantic space will be relatively high. For example, if this interesting story affects the lead actor's public image and echoes other relevant information, then the server will assign it a higher weight coefficient to reflect its importance in the semantic correlation of this entertainment news.
[0023] Step S130, according to the weight coefficients of the first semantic vector units, optimize the first article semantic space to generate x first semantic vector units, and according to the weight coefficients of the second semantic vector units, optimize the second article semantic space to generate x second semantic vector units.
[0024] Still taking the movie review article as an example, assume that the set semantic correlation threshold is an empirical value obtained by the server from processing a large number of movie review articles in the past. If x = 5 is determined according to this threshold (5 here is just for convenience of example). The server sorts the semantic vector units in the first article semantic space according to the weight coefficients of the first semantic vector units in descending order of the weight coefficients (that is, the higher the weight coefficient, the more in the front). Then the first 5 first semantic vector units are selected. For example, in the movie review article, the first 5 semantic vector units with the highest weight coefficients may be semantic vector units related to the core plot of the movie, the wonderful performance of the protagonist, the unique theme expression of the movie, the impressive pictures, and the profound emotions conveyed by the movie, etc.
[0025] For entertainment news articles, follow the same method. If it is determined that x = 5 based on the set semantic correlation threshold and the semantic space dimension of the entertainment news article, arrange the second semantic vector units in descending order according to their weight coefficients, and select the top 5 second semantic vector units. They may be semantic vector units related to important activities of the lead actor, special contributions of the lead actor in the movie, relationships between the lead actor and other actors, highlight images of the lead actor, and unique events related to the movie, etc.
[0026] Step S140: Perform semantic matching analysis on the x first semantic vector units and the x second semantic vector units to generate a first feature identifier corresponding to the x first semantic vector units and a second feature identifier corresponding to the x second semantic vector units. The first feature identifier is used to locate y first semantic vector units among the x first semantic vector units that are semantically relevant to the x second semantic vector units, and the second feature identifier is used to locate z second semantic vector units among the x second semantic vector units that are semantically relevant to the x first semantic vector units. Herein, the first feature identifier and the second feature identifier are used to represent corresponding masks.
[0027] Taking the 5 first semantic vector units in the previously selected movie review article and the 5 second semantic vector units in the entertainment news article as an example. First, determine the first basic feature identifier corresponding to the 5 first semantic vector units and the second basic feature identifier corresponding to the 5 second semantic vector units. For the first basic feature identifier in the movie review article, assume that some words are supplemented during the previous dimension conversion in the movie review article (such as supplementing some professional terms in movie production to better understand the movie review content). Then, in the first basic feature identifier, the markers corresponding to the semantic vectors of these supplementary words are marked as non-core content. For example, if the word "montage technique" is supplemented, the semantic vector corresponding to it in the first basic feature identifier is marked as non-core content.
[0028] Then, screen the 5 first semantic vector units according to the first basic feature identifier to generate 5 first semantic vector units to be optimized. For example, perform special processing on the relevant semantic vector units marked as non-core content to obtain the first semantic vector units to be optimized. Similarly, perform similar operations on the entertainment news article to obtain 5 second semantic vector units to be optimized.
[0029] Next, the server performs semantic relevance analysis on the 5 first semantic vector units to be optimized through the semantic correlation analysis module, and at the same time performs correlation analysis on the 5 second semantic vector units to be optimized and the semantic vector units generated by the semantic correlation analysis module through the interactive correlation analysis module to generate 5 first adjusted semantic vector units. Perform the same operations on the entertainment news article to generate 5 second adjusted semantic vector units.
[0030] For example, in a movie review article, if a first semantic vector unit to be optimized is about a movie scene and it is found through analysis that it has a certain association with the interesting stories about movie shooting shared by the lead actor in entertainment news (the content in the second adjusted semantic vector unit), then this first semantic vector unit to be optimized will be adjusted to a first adjusted semantic vector unit.
[0031] After that, semantic matching analysis is performed on the 5 first adjusted semantic vector units to generate a confidence combination corresponding to each first adjusted semantic vector unit. For example, for the first adjusted semantic vector unit about the movie plot, analyze its association with each semantic vector unit in the entertainment news to obtain the confidence of belonging to the semantically relevant content (such as 0.8, indicating a high probability of relevance) and the confidence of not belonging to the semantically relevant content (such as 0.2). Similarly, similar operations are performed on the 5 second adjusted semantic vector units.
[0032] Based on these confidence combinations, determine 5 first semantic identification codes corresponding to the 5 first adjusted semantic vector units and 5 second semantic identification codes corresponding to the 5 second adjusted semantic vector units. For example, for a certain first adjusted semantic vector unit, if the confidence of belonging to the semantically relevant content is the highest, then the first semantic identification code may be the semantic content label of "plot related".
[0033] Perform logical operation processing on the 5 first semantic identification codes to generate 5 first confidences, and perform logical operation processing on the 5 second semantic identification codes to generate 5 second confidences. Then, screen the 5 first confidences according to the first basic feature identifier to generate a first adjusted feature identifier; screen the 5 second confidences according to the second basic feature identifier to generate a second adjusted feature identifier. Finally, determine the first feature identifier according to the first adjusted feature identifier and determine the second feature identifier according to the second adjusted feature identifier.
[0034] Step S150, perform semantic relevance analysis on the y first semantic vector units and the z second semantic vector units to generate a first semantic association summary result between the first newspaper article and the second newspaper article.
[0035] Suppose after the previous steps, in the movie review article (the first newspaper article), y = 3 first semantic vector units with semantic relevance to the entertainment news article (the second newspaper article) are determined, and in the entertainment news article, z = 2 second semantic vector units with semantic relevance to the movie review article are determined.
[0036] For these 3 first semantic vector units (such as those related to movie plots, lead actor performances, and movie themes) and 2 second semantic vector units (such as those related to the lead actor's contributions in the movie and interesting filming anecdotes shared by the lead actor), the server performs semantic relevance analysis. It analyzes the degree of semantic connection, logical relationships, etc. between these semantic vector units.
[0037] For example, a certain plot in the movie plot is directly related to the interesting filming anecdotes shared by the lead actor, and there is also an internal connection between the lead actor's performance and the lead actor's contributions in the movie. By comprehensively analyzing the relationships between these semantic vector units, the server generates a first semantic association summary result between the first newspaper article (movie review article) and the second newspaper article (entertainment news article). This summary result may be a description indicating how the plot, performance, and theme in the movie review article are related to the lead actor's contributions and interesting filming anecdotes in the entertainment news article, such as "The wonderful plot mentioned in the movie review is a manifestation of the lead actor's contributions in the movie, the wonderful performance of the lead actor echoes the interesting filming anecdotes shared by the lead actor, and the theme of the movie is related to the lead actor's role positioning in the movie" and other similar semantic association descriptions.
[0038] Based on the above steps, the embodiment of the present application significantly improves the accuracy and efficiency of semantic association mining between newspaper articles through semantic extraction and matching analysis. Specifically, first, semantic extraction is performed on the two articles respectively to construct their respective semantic spaces, and through weight coefficient analysis, semantic vector units that have a significant impact on the semantic association of the articles are effectively identified. Subsequently, through semantic space optimization, the most representative semantic vector units are further refined, and the feature identification technology is used to accurately locate the vector units with semantic relevance between the two articles. Finally, through in-depth semantic relevance analysis, a summary result that accurately reflects the semantic association relationship between the two articles is generated, which not only improves the accuracy and comprehensiveness of summary generation, but also greatly promotes the application efficiency in fields such as information retrieval, content recommendation, and public opinion analysis.
[0039] In a possible implementation manner, step S130 includes:
[0040] Step S131, determining x according to a set semantic association degree threshold.
[0041] Step S132, according to the weight coefficients of the respective first semantic vector units, based on the descending order of the weight coefficients, output the first x first semantic vector units in the first article semantic space as the x first semantic vector units, and according to the weight coefficients of the respective second semantic vector units, based on the descending order of the weight coefficients, output the first x second semantic vector units in the second article semantic space as the x second semantic vector units.
[0042] In a possible implementation manner, step S131 includes: determining x according to the semantic correlation degree threshold and the dimension of the first article semantic space, or determining x according to the semantic correlation degree threshold and the dimension of the second article semantic space.
[0043] In this embodiment, still taking the movie review article mentioned above as the first newspaper article and the entertainment news article as the second newspaper article as an example.
[0044] For the process of determining x according to the set semantic correlation degree threshold, when the server processes the movie review article (the first newspaper article) and the entertainment news article (the second newspaper article), there is a set of established standards and processes. The semantic correlation degree threshold is a value preset by the server based on a large amount of text processing experience and the requirement for semantic analysis accuracy. This value reflects the lower limit of the correlation degree that needs to be achieved between semantic vector units when optimizing the semantic space.
[0045] Suppose the set semantic correlation degree threshold is a relative value, for example, 0.6 (this is just an example value set for easy understanding). When processing the movie review article, the server first needs to determine the dimension of the first article semantic space. The dimension of the first article semantic space depends on the complexity and diversity of the semantic information in the movie review article. For example, the movie review article evaluates in detail from multiple aspects such as plot, performance, picture, music, theme, etc. Each aspect contains multiple semantic elements, and these semantic elements together constitute a multi-dimensional semantic space. Suppose after analysis, the dimension of this first article semantic space is 10 (this is just a hypothetical dimension number, and in practice, different values may be obtained according to specific semantic analysis).
[0046] The server determines x according to this semantic correlation degree threshold and the dimension of the first article semantic space. The specific calculation method may be a complex functional relationship. For example, it may be based on the proportional relationship between the semantic correlation degree threshold and the dimension, and through calculation, the number of semantic vector units that meet this threshold requirement is obtained. Suppose according to this calculation method, it is calculated that x = 4. This x = 4 means that when optimizing the first article semantic space, 4 first semantic vector units will be selected.
[0047] A similar process applies to entertainment news articles (the second newspaper articles). Entertainment news articles may contain semantic information in multiple aspects such as the personal lives of the lead actors, the development of their acting careers, and anecdotes related to the movies. This information constitutes the semantic space of the second article. Suppose the dimension of this second article semantic space is 8. Similarly, based on the pre-set semantic correlation threshold of 0.6, through a calculation method similar to that for processing the first article semantic space, the value of x is obtained. Here, it is assumed that the calculated x = 3 (since the dimension of the second article semantic space is different from that of the first article semantic space, the calculated value of x may also be different).
[0048] After determining the value of x, it is necessary to output the first x first semantic vector units in the first article semantic space according to the weight coefficients of each first semantic vector unit, based on the descending order of the weight coefficients. In movie review articles, the weight coefficients of each first semantic vector unit reflect the importance of each semantic vector unit in the semantic association of the entire movie review article. For example, the semantic vector unit related to the movie plot may be given a higher weight coefficient because it is the core content of the movie review. For instance, the semantic vector unit related to the plot "The protagonist encounters a mysterious magical creature during the adventure" has an important impact on the elaboration of the theme and the shaping of the character image in the entire movie review, so the weight coefficient is relatively high. While the semantic vector unit related to a minor detail of a supporting role in the movie has a relatively low weight coefficient because its impact on the semantic association of the overall movie review is relatively small.
[0049] The server sorts all the first semantic vector units in the first article semantic space in descending order according to the weight coefficients. Then, the first x (here x = 4) first semantic vector units are selected as the optimized x first semantic vector units. Suppose the first 4 first semantic vector units ranked are related to the core plot of the movie, the excellent performance of the protagonist, the unique theme expression of the movie, and the impressive scenes respectively.
[0050] Similarly, for entertainment news articles (the second newspaper articles), according to the weight coefficients of each second semantic vector unit, based on the descending order of the weight coefficients, the first x second semantic vector units in the second article semantic space are output as x second semantic vector units. In entertainment news articles, for example, the semantic vector unit regarding a major show business event of the lead actor (such as winning an important award) has a relatively high weight coefficient because it has a greater impact on the semantic association of the entire entertainment news article; while the semantic vector unit regarding a small episode at the event site of the lead actor (such as the interaction details with a small fan) has a relatively low weight coefficient. The server sorts all the second semantic vector units in descending order of the weight coefficient and selects the first x (here x = 3) second semantic vector units as the optimized x second semantic vector units. Suppose these 3 second semantic vector units are respectively related to the important show business event of the lead actor, the special contribution of the lead actor in the movie, and the interesting filming anecdotes shared by the lead actor.
[0051] In this way, the process of determining x according to the semantic association degree threshold and optimizing the first article semantic space and the second article semantic space based on the weight coefficients to generate x first semantic vector units and x second semantic vector units respectively is completed. This process helps to focus on the parts of the article with the closest semantic association, laying a foundation for more accurate semantic analysis and the mining of semantic relationships between the two articles in the future.
[0052] In a possible implementation manner, step S140 includes:
[0053] Step S141, determining the first basic feature identifiers corresponding to the x first semantic vector units, and determining the second basic feature identifiers corresponding to the x second semantic vector units.
[0054] Step S142, screening the x first semantic vector units according to the first basic feature identifiers to generate x first semantic vector units to be optimized, and screening the x second semantic vector units according to the second basic feature identifiers to generate x second semantic vector units to be optimized.
[0055] Step S143, performing semantic relevance analysis on the x first semantic vector units to be optimized by the first semantic association analysis module, and generating x first adjusted semantic vector units according to the first interaction association analysis module for the x second semantic vector units to be optimized and the x semantic vector units generated by the first semantic association analysis module, and performing semantic relevance analysis on the x second semantic vector units to be optimized by the second semantic association analysis module, and generating x second adjusted semantic vector units according to the second interaction association analysis module for the x first semantic vector units to be optimized and the x semantic vector units generated by the second semantic association analysis module.
[0056] Step S144: Adjust the first basic feature identifier according to the x first adjustment semantic vector units to generate a first adjusted feature identifier, and adjust the second basic feature identifier according to the x second adjustment semantic vector units to generate a second adjusted feature identifier.
[0057] Step S145: Determine the first feature identifier according to the first adjusted feature identifier, and determine the second feature identifier according to the second adjusted feature identifier.
[0058] In this embodiment, taking the previously mentioned movie review article (the first newspaper article) and entertainment news article (the second newspaper article) as examples, the movie review article obtains x first semantic vector units through the previous steps, and the entertainment news article also obtains x second semantic vector units.
[0059] First, determine the first basic feature identifier corresponding to the x first semantic vector units, and determine the second basic feature identifier corresponding to the x second semantic vector units. For the x first semantic vector units in the movie review article, assume that these x first semantic vector units are semantic vector units related to the core plot of the movie, the wonderful performance of the protagonist, the unique theme expression of the movie, and the impressive pictures respectively. When determining the first basic feature identifier, the server will consider some preprocessing operations that may have been performed on the previous movie review article or factors such as the structure of the article itself. For example, if some vocabulary about photography techniques is added during the semantic extraction of the movie review article to better understand the semantic vector units related to the movie pictures, then in the first basic feature identifier, the part of the semantic vector units related to these added vocabulary will be marked as auxiliary content or non-core content. For example, for the semantic vector units related to the pictures, if the vocabulary "use of long shots" is added, then the corresponding semantic vector part of "use of long shots" in the first basic feature identifier will be marked as non-core content to indicate its relatively secondary position in the entire semantic association.
[0060] For the x second semantic vector units in the entertainment news article, assume that these x second semantic vector units are semantic vector units related to the important acting activities of the leading actor, the special contributions of the leading actor in the movie, and the interesting filming anecdotes shared by the leading actor respectively. Similar factors will also be considered when determining the second basic feature identifier. If some vocabulary is added to supplement the acting background information of the leading actor during the processing of the entertainment news article, such as "learning experience in the drama school", then in the second basic feature identifier, the part of the semantic vector units corresponding to these added vocabulary will be marked as non-core content.
[0061] Next, x first semantic vector units are screened according to the first basic feature identifier to generate x first semantic vector units to be optimized, and x second semantic vector units are screened according to the second basic feature identifier to generate x second semantic vector units to be optimized. For movie review articles, the server selects the first semantic vector units marked as non-core content or having more auxiliary content according to the marks in the first basic feature identifier for special processing to form x first semantic vector units to be optimized. For example, if there are more non-core content marks corresponding to supplementary words in the semantic vector units related to the picture when determining the first basic feature identifier before, then this semantic vector unit related to the picture may be selected into the first semantic vector units to be optimized.
[0062] For entertainment news articles, screening is also carried out according to the second basic feature identifier. If there are more non-core content marks corresponding to the supplementary acting background information in the semantic vector units related to the main actor's performing activities, then this semantic vector unit may be selected into the x second semantic vector units to be optimized.
[0063] Then, the first semantic association analysis module performs semantic association analysis on the x first semantic vector units to be optimized, and the first interaction association analysis module generates x first adjusted semantic vector units for the x second semantic vector units to be optimized and the x semantic vector units generated by the first semantic association analysis module. Also, the second semantic association analysis module performs semantic association analysis on the x second semantic vector units to be optimized, and the second interaction association analysis module generates x second adjusted semantic vector units for the x first semantic vector units to be optimized and the x semantic vector units generated by the second semantic association analysis module.
[0064] In the case of movie review articles, the first semantic association analysis module deeply analyzes the semantic associations among the x first semantic vector units to be optimized. For example, for the first semantic vector units to be optimized related to the movie plot and the protagonist's performance, analyze their logical connections and semantic association degrees in the movie review article. At the same time, the first interaction association analysis module considers the interaction relationship between the x second semantic vector units (from entertainment news articles) and the semantic vector units generated by the first semantic association analysis module. Suppose the first semantic association analysis module discovers a potential association between the first semantic vector unit to be optimized related to the movie plot and the interesting filming anecdotes shared by the main actor in the entertainment news article (content in the second semantic vector unit to be optimized) during the analysis. Then, through further analysis and processing by the first interaction association analysis module, the first semantic vector unit to be optimized related to the movie plot will be adjusted to more accurately reflect the semantic association with the entertainment news article, thereby generating one of the x first adjusted semantic vector units.
[0065] For entertainment news articles, the second semantic correlation analysis module performs semantic correlation analysis on x second semantic vector units to be optimized. For example, it analyzes the correlation between two second semantic vector units to be optimized, such as the important performing activities of the lead actor and the special contributions of the lead actor in the movie. The second interaction correlation analysis module then considers the interaction relationship between x first semantic vector units to be optimized (from movie review articles) and the semantic vector units generated by the second semantic correlation analysis module. For instance, when analyzing the second semantic vector unit to be optimized regarding the special contributions of the lead actor in the movie, it is found that there is a correlation with the wonderful performance of the protagonist in the movie review article (content in the first semantic vector unit to be optimized). Through the processing of the second interaction correlation analysis module, the semantic vector unit regarding the special contributions of the lead actor in the movie is adjusted, and thus one of the x second adjusted semantic vector units is generated.
[0066] After that, based on the x first adjusted semantic vector units, the first basic feature identifier is adjusted to generate the first adjusted feature identifier, and based on the x second adjusted semantic vector units, the second basic feature identifier is adjusted to generate the second adjusted feature identifier. For movie review articles, the server will adjust the first basic feature identifier according to the new semantic relationships and content of the x first adjusted semantic vector units. For example, if a certain first adjusted semantic vector unit (such as related to the movie plot) has a stronger correlation with the semantic vector units in the entertainment news article after the previous adjustment, then in the first adjusted feature identifier, the part originally marked as non-core content may be adjusted according to the new semantic correlation situation. If a certain auxiliary word related to the plot (such as the specific name of a certain scene) becomes more important under the new semantic correlation, then the marking of the semantic vector part corresponding to this word in the first adjusted feature identifier will be adjusted from non-core content to relatively core content.
[0067] For entertainment news articles, the second basic feature identifier is also adjusted based on the x second adjusted semantic vector units. For example, after the adjustment of the second adjusted semantic vector unit regarding the special contributions of the lead actor in the movie, if a new and closer semantic connection is established with the semantic vector units in the movie review article, then in the second adjusted feature identifier, the marked content related to this semantic vector unit (such as some specific details describing the actor's contributions originally marked as non-core content) may be adjusted according to the new semantic correlation situation.
[0068] Finally, determine the first feature identifier based on the first adjustment feature identifier, and determine the second feature identifier based on the second adjustment feature identifier. The first feature identifier is further determined on the basis of the first adjustment feature identifier. It synthesizes all the previous adjustment and analysis results regarding the first semantic vector unit and can accurately locate the part of the x first semantic vector units that has semantic relevance to the x second semantic vector units. For example, in a movie review article, the first feature identifier can clearly indicate which first semantic vector units (such as semantic vector units related to movie plots, lead actor performances, etc.) have strong semantic relevance to the semantic vector units in an entertainment news article.
[0069] Similarly, the second feature identifier is determined based on the second adjustment feature identifier, and it can accurately locate the part of the x second semantic vector units that has semantic relevance to the x first semantic vector units. For example, in an entertainment news article, the second feature identifier can point out the semantic relevance between semantic vector units such as important acting activities of the lead actor and the lead actor's special contributions in the movie and the semantic vector units in the movie review article. Through such a process, the server completes the operation of semantic matching analysis on the x first semantic vector units and the x second semantic vector units, generating the corresponding first feature identifier and second feature identifier, which helps to further explore the semantic association relationship between the two articles.
[0070] In a possible implementation manner, before step S110, when the dimensions of the first newspaper article and the second newspaper article are inconsistent, the first newspaper article and the second newspaper article can be converted to the same dimension.
[0071] Step S141 includes:
[0072] Step S1411, determine the first basic feature identifier according to the words supplemented in the dimension conversion of the first newspaper article. In the first basic feature identifier, the markers corresponding to the semantic vectors of the words supplemented in the x first semantic vector units are marked as non-core content.
[0073] Step S1412, determine the second basic feature identifier according to the words supplemented in the dimension conversion of the second newspaper article. In the second basic feature identifier, the markers corresponding to the semantic vectors of the words supplemented in the x second semantic vector units are marked as non-core content.
[0074] In this embodiment, still taking the movie review article mentioned above as the first newspaper article and the entertainment news article as the second newspaper article as an example.
[0075] Before generating the first article semantic space by semantic extraction of movie review articles and the second article semantic space by semantic extraction of entertainment news articles, the server needs to first handle the case of inconsistent dimensions between the first newspaper article (movie review article) and the second newspaper article (entertainment news article). Due to differences in the nature of the content, the scope covered, and the focus of expression between movie review articles and entertainment news articles, their semantic dimensions may vary.
[0076] For example, movie review articles may focus on the artistic aspects of movies, constructing semantic information from multiple artistic elements such as plot, acting, visuals, and music. These elements are interrelated, forming the unique semantic dimension of movie review articles. Entertainment news articles, on the other hand, pay more attention to artist-related events, such as the activities of the lead actors, image building, public reactions, etc., which constitute the semantic dimension of entertainment news articles. It is very likely that the semantic dimension of movie review articles is composed of multiple sub-dimensions related to movie art, while the semantic dimension of entertainment news articles is composed of multiple sub-dimensions related to artist activities. The number and specific content of their sub-dimensions may be different, resulting in inconsistent overall dimensions.
[0077] When the server detects such a dimension inconsistency, it needs to convert the movie review articles and entertainment news articles to the same dimension. This conversion process is not simply about unifying the dimension values, but rather through a method of semantic mapping and supplementation. For movie review articles, if their semantic expression in a certain artistic element is relatively in-depth, while the entertainment news articles lack the corresponding semantic dimension in this regard, the server may supplement some relevant general semantic elements in the semantic construction of entertainment news articles to align their dimensions with those of movie review articles to a certain extent. Vice versa.
[0078] Suppose the movie review article has a very detailed description of the visual analysis, covering semantic information in multiple dimensions such as color matching and camera language, while the entertainment news article hardly mentions this aspect. To convert the two to the same dimension, the server will supplement some basic semantic elements about movie visuals in the semantic construction of the entertainment news article, such as simply mentioning whether the visual effect of the picture is attractive. Similarly, if the entertainment news article has a detailed description of the fan base of the lead actor, while the movie review article does not cover this content, the server will supplement some semantic elements about audience reactions (a broad concept similar to the reactions of the fan base) in the semantic construction of the movie review article.
[0079] After completing such dimensionality conversion, start determining the first basic feature identifiers corresponding to the x first semantic vector units and the second basic feature identifiers corresponding to the x second semantic vector units. For the x first semantic vector units in the movie review article, since some words are supplemented during the dimensionality conversion process, these supplemented words have special significance when determining the first basic feature identifiers. For example, when converting the movie review article and the entertainment news article to the same dimension, in order to make the semantics of the movie review article more comprehensively aligned with the entertainment news article, some words about the market response of the movie are supplemented in the movie review article (although the movie review itself may rarely directly mention this aspect).
[0080] When determining the first basic feature identifiers, for the x first semantic vector units, if a certain semantic vector unit has a semantic association with the supplemented words about the market response of the movie, then in the first basic feature identifier, the part corresponding to the semantic vector of the supplemented word in this semantic vector unit will be marked as non-core content. For example, in the movie review article, there is a first semantic vector unit about the expression of the movie theme. During the dimensionality conversion process, in order to align with the dimension of the movie's influence in the entertainment news article, some words about the box office performance of the movie are supplemented. This box office performance word is associated with the semantic vector unit of the movie theme expression. However, in the semantic system of the entire movie review article, compared with the in-depth elaboration of the movie theme itself, the relationship with the plot and performance, etc., the box office performance belongs to relatively non-core content. Therefore, in the first basic feature identifier, the semantic vector part corresponding to this supplemented word of box office performance will be marked as non-core content.
[0081] A similar situation applies to the x second semantic vector units in the entertainment news article. Suppose that when converting the entertainment news article and the movie review article to the same dimension, some basic information words about the movie production team are supplemented in the entertainment news article (because the movie review article may involve some indirect evaluations of the production team, such as the influence of the director's style on the movie, etc.). When determining the second basic feature identifiers, for the x second semantic vector units, if a certain semantic vector unit has a semantic association with the supplemented words about the movie production team, then in the second basic feature identifier, the part corresponding to the semantic vector of the supplemented word in this semantic vector unit will be marked as non-core content. For example, in the entertainment news article, there is a second semantic vector unit about the character shaping of the lead actor in the movie. After supplementing the words related to the movie production team, a certain member of the production team (such as the producer) may have a certain influence on the character shaping. However, compared with the core content such as the lead actor's own acting skills and the character's position in the plot, this part related to the producer will be marked as non-core content in the second basic feature identifier.
[0082] In this way, after processing the inconsistent dimensions of the first newspaper article and the second newspaper article and performing conversion, the first basic feature identifier and the second basic feature identifier are accurately determined according to the supplemented vocabulary in the dimension conversion, laying a foundation for further analyzing the semantic relationship between the two articles. This way of marking non-core content helps to more specifically focus on the core semantic part in subsequent semantic analysis, improving the accuracy of semantic matching and correlation analysis.
[0083] In a possible implementation manner, step S144 includes:
[0084] Step S1441: Perform semantic matching analysis on the x first adjusted semantic vector units to generate a confidence combination corresponding to each first adjusted semantic vector unit among the x first adjusted semantic vector units, and perform semantic matching analysis on the x second adjusted semantic vector units to generate a confidence combination corresponding to each second adjusted semantic vector unit among the x second adjusted semantic vector units. The confidence combination reflects the confidence of belonging to semantically related content and the confidence of not belonging to semantically related content. According to the confidence combination corresponding to each first adjusted semantic vector unit, determine the x first semantic identification codes corresponding to the x first adjusted semantic vector units, and according to the confidence combination corresponding to each second adjusted semantic vector unit, determine the x second semantic identification codes corresponding to the x second adjusted semantic vector units. The first semantic identification code reflects the semantic content label corresponding to the maximum confidence in the confidence combination corresponding to the first adjusted semantic vector unit, and the second semantic identification code reflects the semantic content label corresponding to the maximum confidence in the confidence combination corresponding to the second adjusted semantic vector unit. Perform logical operation processing on the x first semantic identification codes to generate x first confidences, and perform logical operation processing on the x second semantic identification codes to generate x second confidences. Screen the x first confidences according to the first basic feature identifier to generate the first adjusted feature identifier, and screen the x second confidences according to the second basic feature identifier to generate the second adjusted feature identifier.
[0085] Continuing with the previously mentioned movie review article as the first newspaper article and the entertainment news article as the second newspaper article as an example.
[0086] In this process, the server first performs semantic matching analysis on the x first adjusted semantic vector units to generate a confidence combination corresponding to each first adjusted semantic vector unit among the x first adjusted semantic vector units, and at the same time performs semantic matching analysis on the x second adjusted semantic vector units to generate a confidence combination corresponding to each second adjusted semantic vector unit among the x second adjusted semantic vector units.
[0087] For the x first adjusted semantic vector units in the movie review article, take the first adjusted semantic vector unit regarding the movie plot as an example. The server will conduct a detailed semantic matching analysis of this vector unit with each semantic element in the entertainment news article. For example, there may be a connection between a certain key plot in the movie plot and the interesting filming anecdotes shared by the lead actor in the entertainment news article. The server will use a series of semantic analysis algorithms to determine the possibility of this connection, thereby generating the confidence level of semantic-related content and the confidence level of non-semantic-related content. Suppose after analysis, the confidence level of the connection between this movie plot and the lead actor's filming anecdotes is 0.7 (indicating a relatively high possibility of relevance), and the confidence level of non-semantic-related content is 0.3. For the first adjusted semantic vector unit related to the movie theme, after semantic matching analysis with the lead actor's role positioning in the movie and other content in the entertainment news article, the confidence level of semantic-related content is 0.6, and the confidence level of non-semantic-related content is 0.4, etc., thus generating the corresponding confidence level combinations for each first adjusted semantic vector unit.
[0088] A similar operation is carried out for the x second adjusted semantic vector units in the entertainment news article. For example, for the second adjusted semantic vector unit of the lead actor's important acting activities, the server will conduct a semantic matching analysis of it with each semantic element in the movie review article. If the acting activity of the lead actor is matched and analyzed with the content related to the promotional effect of the movie in the movie review article, the confidence level of semantic-related content is obtained as 0.8, and the confidence level of non-semantic-related content is 0.2. For the second adjusted semantic vector unit of the lead actor's special contribution in the movie, after semantic matching analysis with the content related to the artistic achievements of the movie in the movie review article, the confidence level of semantic-related content is 0.7, and the confidence level of non-semantic-related content is 0.3, etc., thereby generating the corresponding confidence level combinations for each second adjusted semantic vector unit.
[0089] Next, based on the confidence combinations corresponding to each first adjusted semantic vector unit, x first semantic identification codes corresponding to the x first adjusted semantic vector units are determined, and based on the confidence combinations corresponding to each second adjusted semantic vector unit, x second semantic identification codes corresponding to the x second adjusted semantic vector units are determined. For the first adjusted semantic vector unit in the movie review article, taking the confidence combination corresponding to the previous movie plot (the confidence of belonging to the semantically related content is 0.7, and the confidence of not belonging to the semantically related content is 0.3) as an example, since the confidence of belonging to the semantically related content is relatively high, the corresponding first semantic identification code may be a semantic content label such as "plot - related to entertainment news", indicating that this first adjusted semantic vector unit is mainly identified by the characteristics related to the plot and entertainment news in the semantic association with the entertainment news article. For the first adjusted semantic vector unit related to the movie theme, the confidence of belonging to the semantically related content in its confidence combination is 0.6, and the confidence of not belonging to the semantically related content is 0.4, and the corresponding first semantic identification code may be "theme - partially related to entertainment news".
[0090] For the second adjusted semantic vector unit in the entertainment news article, the confidence combination corresponding to the important acting activities of the lead actor (the confidence of belonging to the semantically related content is 0.8, and the confidence of not belonging to the semantically related content is 0.2), and the corresponding second semantic identification code may be "acting activity - strongly related to movie review". The confidence combination corresponding to the special contribution of the lead actor in the movie (the confidence of belonging to the semantically related content is 0.7, and the confidence of not belonging to the semantically related content is 0.3), and the corresponding second semantic identification code may be "special contribution - relatively related to movie review".
[0091] Then, logical operation processing is performed on the x first semantic identification codes to generate x first confidences, and logical operation processing is performed on the x second semantic identification codes to generate x second confidences. Suppose for the x first semantic identification codes in the movie review article, the server adopts a weighted summation logical operation method (this is just an example method). If the first semantic identification code "plot - related to entertainment news" is given a higher weight and "theme - partially related to entertainment news" is given a lower weight, a value is obtained through weighted summation calculation as one of the x first confidences. For the x second semantic identification codes in the entertainment news article, a similar logical operation method is also adopted. For example, the second semantic identification code "acting activity - strongly related to movie review" corresponding to the important acting activities of the lead actor has a higher weight, and the second semantic identification code "special contribution - relatively related to movie review" corresponding to the special contribution of the lead actor in the movie has a lower weight, and a value is obtained through weighted summation calculation as one of the x second confidences.
[0092] After that, x first confidence levels are selected based on the first basic feature identifier to generate a first adjusted feature identifier, and x second confidence levels are selected based on the second basic feature identifier to generate a second adjusted feature identifier. For movie review articles, the first basic feature identifier marks which parts are core content and which parts are non-core content. Suppose in the first basic feature identifier, the parts related to the movie plot are marked as core content, and the parts related to the movie's market response (content supplemented in the previous dimensional transformation) are marked as non-core content. When selecting x first confidence levels, if the first semantic identification code corresponding to a certain first confidence level is related to the plot and has a relatively high weight in the logical operation, then this first confidence level will be given key consideration when generating the first adjusted feature identifier; if the first semantic identification code corresponding to a certain first confidence level is related to the movie's market response and has a relatively low weight in the logical operation, then its influence on the first adjusted feature identifier will be relatively small when screened according to the first basic feature identifier.
[0093] For entertainment news articles, the second basic feature identifier marks the core and non-core parts of the content related to the lead actor. For example, the content related to the acting skills of the lead actor is marked as core content, and the tabloid news of the lead actor (related content supplemented in the previous dimensional transformation) is marked as non-core content. When selecting x second confidence levels, if the second semantic identification code corresponding to a certain second confidence level is related to the acting skills and has a relatively high weight in the logical operation, then this second confidence level will be given key consideration when generating the second adjusted feature identifier; if the second semantic identification code corresponding to a certain second confidence level is related to the tabloid news and has a relatively low weight in the logical operation, then its influence on the second adjusted feature identifier will be relatively small when screened according to the second basic feature identifier.
[0094] Alternatively, in step S1442, based on the x first adjustment semantic vector units, the x second adjustment semantic vector units, the first adjustment feature identifier, and the second adjustment feature identifier, adjust the x first adjustment semantic vector units to generate x third adjustment semantic vector units, and based on the x first adjustment semantic vector units, the x second adjustment semantic vector units, the first adjustment feature identifier, and the second adjustment feature identifier, adjust the x second adjustment semantic vector units to generate x fourth adjustment semantic vector units. Based on the x third adjustment semantic vector units, adjust the first adjustment feature identifier to generate a third adjustment feature identifier, and based on the x fourth adjustment semantic vector units, adjust the second adjustment feature identifier to generate a fourth adjustment feature identifier. Determine the first feature identifier based on the third adjustment feature identifier, and determine the second feature identifier based on the fourth adjustment feature identifier. The first feature identifier is the adjustment feature identifier obtained when the adjustment cycle times of the first basic feature identifier are not less than the set order, and the second feature identifier is the adjustment feature identifier obtained when the adjustment cycle times of the second basic feature identifier are not less than the set order.
[0095] In another case, based on the x first adjustment semantic vector units, the x second adjustment semantic vector units, the first adjustment feature identifier, and the second adjustment feature identifier, adjust the x first adjustment semantic vector units to generate x third adjustment semantic vector units, and adjust the x second adjustment semantic vector units to generate x fourth adjustment semantic vector units. For the x first adjustment semantic vector units in the movie review article, assuming the first adjustment semantic vector unit related to the movie plot, according to its sharing of the filming anecdotes with the lead actor in the entertainment news article (from the x second adjustment semantic vector units) and the current first adjustment feature identifier (such as marking the importance degree of the plot-related part), if it is found that a certain detail in the plot has a deeper semantic association with the filming anecdotes but has not been fully reflected before, then adjust the first adjustment semantic vector unit related to the movie plot to generate one of the x third adjustment semantic vector units.
[0096] For the x second adjustment semantic vector units in the entertainment news article, such as the second adjustment semantic vector unit of the special contribution of the lead actor in the movie, according to its connection with the movie theme in the movie review article (from the x first adjustment semantic vector units) and the current second adjustment feature identifier (such as marking the status of the special contribution in the overall semantics), if it is found that there is a new semantic connection between the special contribution and the movie theme that has not been explored, then adjust this semantic vector unit to generate one of the x fourth adjustment semantic vector units.
[0097] Finally, the first adjusted feature identifier is adjusted based on x third adjusted semantic vector units to generate a third adjusted feature identifier, and the second adjusted feature identifier is adjusted based on x fourth adjusted semantic vector units to generate a fourth adjusted feature identifier. For movie review articles, according to the new semantic relationships and content in the x third adjusted semantic vector units, if a certain third adjusted semantic vector unit (such as those related to movie plots) has new important information in the semantic association with entertainment news articles, then in the first adjusted feature identifier, the original markings for the plot-related parts (such as the degree of importance, etc.) will be adjusted according to this new situation to generate the third adjusted feature identifier.
[0098] For entertainment news articles, according to the new semantic relationships and content in the x fourth adjusted semantic vector units, if a certain fourth adjusted semantic vector unit (such as those related to the special contributions of the lead actors in the movie) has new situations in the semantic association with movie review articles, then in the second adjusted feature identifier, the original markings for the special contribution-related parts will be adjusted according to the new situation to generate the fourth adjusted feature identifier. The first feature identifier is determined based on the third adjusted feature identifier, and the second feature identifier is determined based on the fourth adjusted feature identifier. Here, the first feature identifier is an adjusted feature identifier obtained after the first basic feature identifier has been adjusted multiple times (the number of adjustment cycles is not less than the set order), and the second feature identifier is an adjusted feature identifier obtained after the second basic feature identifier has been adjusted multiple times (the number of adjustment cycles is not less than the set order). They can more accurately locate the semantic correlation between the x first semantic vector units and the x second semantic vector units.
[0099] In a possible implementation manner, step S110 includes: performing semantic extraction on the first newspaper article to generate the first article semantic space and the third article semantic space, and performing semantic extraction on the second newspaper article to generate the second article semantic space and the fourth article semantic space. The dimension of the first article semantic space is different from the dimension of the third article semantic space, and the dimension of the second article semantic space is different from the dimension of the fourth article semantic space.
[0100] After step S150, the method further includes:
[0101] Step S160: Optimize the third article semantic space based on the first semantic association summary result, the dimension of the first article semantic space, the dimension of the third article semantic space, and the third article semantic space to generate y newspaper article segments corresponding to the y first semantic vector units, and optimize the fourth article semantic space based on the first semantic association summary result, the dimension of the second article semantic space, the dimension of the fourth article semantic space, and the fourth article semantic space to generate z newspaper article segments corresponding to the z second semantic vector units.
[0102] Step S170: Semantically enhance the y newspaper article segments based on the y newspaper article segments and the z newspaper article segments to generate y enhanced segments, and semantically enhance the z newspaper article segments based on the y newspaper article segments and the z newspaper article segments to generate z enhanced segments.
[0103] Step S180: Extract the core semantics of the y enhanced segments to generate y core semantics, and perform semantic compression or expansion processing on the z enhanced segments based on the semantic depth error between the third article semantic space and the fourth article semantic space to generate z adjusted segments.
[0104] Step S190: Perform semantic localization processing on the z adjusted segments to generate z localization semantics corresponding to the z adjusted segments, and perform semantic association analysis on the y core semantics and the z localization semantics to generate a second semantic association summary result between the first newspaper article and the second newspaper article.
[0105] Among them, Step S190 includes:
[0106] Step S191: Obtain the semantic importance coefficients of each semantic unit in the first adjusted segment among the z adjusted segments.
[0107] Step S192: Perform weighted fusion on the semantic importance coefficients and semantic contents of each semantic unit in the first adjusted segment to generate the localization semantics of the first adjusted segment.
[0108] Still taking the movie review article mentioned above as the first newspaper article and the entertainment news article as the second newspaper article as an example.
[0109] When the server performs semantic extraction on the first newspaper article (movie review article), it will generate the first article semantic space and the third article semantic space. The movie review article contains rich and diverse semantic information, and different semantic spaces with different dimensions will be formed due to different semantic analysis perspectives and requirements.
[0110] For the generation of the semantic space of the first article, the server focuses on the main semantic elements in the film review article, such as evaluations of the plot, acting, visuals, music, etc. It analyzes the semantic relationships between each word and sentence and maps these elements into a multi-dimensional space. For example, the semantics of the plot may form a subspace in this space, which contains semantic vectors such as the beginning, development, climax, and end of the story, the rationality of the plot, and the relevance between the plot and the theme; the acting aspect contains semantic vectors such as the actor's interpretation of the role, the level of acting skills, and the cooperation with other actors. These semantic vectors together constitute the semantic space of the first article, which focuses on the semantic expression of the core content of the film review article.
[0111] The generation of the semantic space of the third article may start from different perspectives. Suppose it pays more attention to the background information in the film review article, the relevant cultural elements cited, or some supplementary content that helps understand the film review. For example, the semantic relationship construction of the relevant plot of the original novel from which the film is adapted (although not the core of the film review but as supplementary information), the cultural background involved in the film (such as the influence of the cultural customs of a specific region on the film plot), etc. In this way, the semantic space of the third article is different from that of the first article in dimension because they cover different semantic focuses and perspectives.
[0112] Similarly, for the second newspaper article (entertainment news article), the server generates the semantic space of the second article and the semantic space of the fourth article during semantic extraction. The semantic space of the second article focuses on the core content in the entertainment news article, such as semantic information about the acting activities of the leading actor, public image, relationship with other artists, etc. For example, the acting activities of the leading actor may include recently participated movies, variety shows attended, etc. These contents form their respective semantic vectors in the semantic space, such as the type of the participated movie, the role positioning in the variety show, etc., and together constitute the semantic space of the second article.
[0113] The semantic space of the fourth article is constructed from different semantic perspectives and may focus on some peripheral information in the entertainment news article, such as the background of the reporting source, the relationship between the news release time and the overall entertainment industry environment at that time. This makes the semantic space of the fourth article different from that of the second article in dimension, and it contains some auxiliary or background semantic elements.
[0114] After generating the first semantic association summary result between the first newspaper article and the second newspaper article, a series of operations are carried out next.
[0115] First, based on the first semantic association summary result, the dimension of the first article semantic space, the dimension of the third article semantic space, and the third article semantic space, semantic space optimization is performed on the third article semantic space to generate y newspaper article segments corresponding to y first semantic vector units. The first semantic association summary result contains the main semantic association information between movie review articles and entertainment news articles. For example, the first semantic association summary result indicates that there is a strong association between the movie plot in the movie review article and the performance of the lead actor in the movie in the entertainment news article. Based on this result, when optimizing the third article semantic space, the server will focus on background information or supplementary content related to the movie plot. If there are cultural background elements related to the movie plot in the third article semantic space (such as cultural elements where the movie plot takes place in a specific historical period), and these elements have a certain potential connection with the entertainment news article in the first semantic association summary result, the server will select these parts to generate y newspaper article segments corresponding to y first semantic vector units.
[0116] Meanwhile, based on the first semantic association summary result, the dimension of the second article semantic space, the dimension of the fourth article semantic space, and the fourth article semantic space, semantic space optimization is performed on the fourth article semantic space to generate z newspaper article segments corresponding to z second semantic vector units. For example, the first semantic association summary result shows that there is a certain association between the image of the lead actor in the entertainment news article and the movie theme in the movie review article. Then, when optimizing the fourth article semantic space, the server will search for background information related to the image of the lead actor (such as the relationship between the image strategy of the lead actor at the time of news release and the overall atmosphere in the entertainment industry at that time), select the relevant parts, and generate z newspaper article segments corresponding to z second semantic vector units.
[0117] Next, semantic enhancement is performed on the y newspaper article segments based on the y newspaper article segments and the z newspaper article segments to generate y enhanced segments, and semantic enhancement is performed on the z newspaper article segments to generate z enhanced segments. For the y newspaper article segments (from the third article semantic space related to movie review articles), assuming that one of the newspaper article segments is cultural background information about the movie plot, the server will use the information in the z newspaper article segments (from the fourth article semantic space related to entertainment news articles) for semantic enhancement. For example, if the lead actor in the entertainment news article mentions their understanding of the cultural background of the movie plot during an interview, the server will incorporate this information into the newspaper article segment about the cultural background information of the movie plot to generate an enhanced segment. Similarly, for one of the z newspaper article segments, such as a segment about the relationship between the lead actor's image strategy and the entertainment industry atmosphere, semantic enhancement is performed using the movie theme-related information in the y newspaper article segments (if the movie theme involves a reflection of a social phenomenon that has a certain connection with the entertainment industry atmosphere at that time) to generate one of the z enhanced segments.
[0118] Then, the core semantics of the y enhanced segments are extracted to generate y core semantics. For each of the y enhanced segments, the server will deeply analyze its semantic structure to find the core part that best represents the semantic essence of this segment. For example, for an enhanced segment about the cultural background information of the movie plot after semantic enhancement, its core semantics may be the influence of a specific cultural background on the behavioral motivation of the characters in the movie plot. At the same time, based on the semantic depth error between the third article semantic space and the fourth article semantic space, semantic compression or expansion processing is performed on the z enhanced segments to generate z adjusted segments. Assuming that the third article semantic space has a deeper semantic depth related to the movie plot than the fourth article semantic space in terms of the lead actor's image (i.e., there is a semantic depth error), for one of the z enhanced segments about the relationship between the lead actor's image strategy and the entertainment industry atmosphere, if the semantic information in this segment is relatively shallow, the server may perform expansion processing according to the semantic depth error to supplement more information related to the lead actor's image and related to the semantics of the movie review article to generate one of the z adjusted segments; conversely, if the semantic information in this segment is too complex and needs to be simplified according to the semantic depth error, semantic compression processing is performed.
[0119] After that, semantic localization processing is performed on z adjustment segments to generate z localization semantics corresponding to the z adjustment segments. Taking the first adjustment segment among the z adjustment segments as an example, the server first obtains the semantic importance coefficients of each semantic unit in the first adjustment segment. For example, in a first adjustment segment regarding the adjustment of the leading actor's image strategy, the semantic unit of the change in the leading actor's image positioning may have a relatively high semantic importance coefficient because it has an important impact on the semantic core of the entire segment; while the semantic unit of the news source reporting this change in image positioning may have a relatively low semantic importance coefficient. Then, the server performs weighted fusion on the semantic importance coefficients and semantic contents of each semantic unit in the first adjustment segment to generate the localization semantics of the first adjustment segment. For example, multiply the semantic content of the semantic unit of the change in the leading actor's image positioning by its relatively high semantic importance coefficient, and then add the semantic content of the semantic unit of the news source multiplied by its relatively low semantic importance coefficient. The result obtained after such weighted fusion is the localization semantics of the first adjustment segment.
[0120] Finally, semantic association analysis is performed on the y core semantics and the z localization semantics to generate a second semantic association summary result between the first newspaper article and the second newspaper article. For example, there may be a certain logical connection between the influence of the cultural background of the movie plot on the character's behavioral motivation among the y core semantics and the change in the leading actor's image positioning among the z localization semantics. It may be that the cultural background in the movie plot affects the audience's expectations for the leading actor's image, thus prompting the leading actor to adjust the image positioning. By performing such a comprehensive semantic association analysis on all y core semantics and z localization semantics, the server generates a more in-depth and comprehensive second semantic association summary result, which further reveals the deep semantic connections between the movie review article and the entertainment news article, including not only the superficial associations but also the semantic relationships hidden behind the articles discovered through operations such as semantic space optimization, semantic enhancement, and semantic adjustment.
[0121] In a possible implementation manner, the method further includes:
[0122] Step A110, performing semantic extraction on the first template newspaper article to generate a first template article semantic space, and performing semantic extraction on the second template newspaper article to generate a second template article semantic space.
[0123] Step A120, based on a deep learning network, performing semantic matching analysis on the first template article semantic space to generate the weight coefficients of each first semantic vector unit in the first template article semantic space, and based on the deep learning network, performing semantic matching analysis on the second template article semantic space to generate the weight coefficients of each second semantic vector unit in the second template article semantic space. The weight coefficients reflect the influence coefficients on the semantic association of the template newspaper article.
[0124] Step A130: Optimize the semantic space of the first template article semantic space according to the weight coefficients of the first semantic vector units, to generate x first semantic vector units, and optimize the semantic space of the second template article semantic space according to the weight coefficients of the second semantic vector units, to generate x second semantic vector units.
[0125] Step A140: Perform semantic matching analysis on the x first semantic vector units and the x second semantic vector units according to the semantic matching analysis network, to generate a first feature identifier corresponding to the x first semantic vector units and a second feature identifier corresponding to the x second semantic vector units. The first feature identifier is used to locate y first semantic vector units among the x first semantic vector units that are semantically relevant to the x second semantic vector units, and the second feature identifier is used to locate z second semantic vector units among the x second semantic vector units that are semantically relevant to the x first semantic vector units.
[0126] Step A150: Optimize the deep learning network according to the weight coefficients of the first semantic vector units and the weight coefficients of the second semantic vector units, and optimize the semantic matching analysis network according to the first feature identifier and the second feature identifier.
[0127] In a possible implementation manner, step A150 includes:
[0128] Step A151: Determine a first training cost according to the error between the feature identifier reflecting whether there is a topic feature vector in the first semantic vector units in the first template article semantic space and the weight coefficients of the first semantic vector units in the first template article semantic space. And determine a second training cost according to the error between the feature identifier reflecting whether there is a topic feature vector in the second semantic vector units in the second template article semantic space and the weight coefficients of the second semantic vector units in the second template article semantic space.
[0129] Step A152: Optimize the deep learning network according to the first training cost and the second training cost.
[0130] Step A153: Determine a third training cost according to the error between the first feature identifier and the semantic correlation feature identifier, and determine a fourth training cost according to the error between the second feature identifier and the semantic correlation feature identifier. Wherein, the semantic correlation feature identifier is a feature identifier reflecting whether there is a topic feature vector in the semantically relevant part of the x first semantic vector units and the x second semantic vector units.
[0131] Step A154, optimize the semantic matching analysis network according to the third training cost and the fourth training cost.
[0132] In this embodiment, continue to use the previous example related to the movie review article and the entertainment news article to analogize the first template newspaper article and the second template newspaper article. Assume that the first template newspaper article is a retrospective movie review template article about a classic movie series, and the second template newspaper article is a retrospective entertainment news template article about the acting career of the lead actor in this movie series.
[0133] First, the server performs semantic extraction on the first template newspaper article to generate a first template article semantic space, and performs semantic extraction on the second template newspaper article to generate a second template article semantic space.
[0134] For the first template newspaper article (retrospective movie review template article about a classic movie series), when the server performs semantic extraction, it will perform semantic analysis on various elements in the article. For example, for each movie in the movie series, it will analyze its respective plot features, the overall theme evolution of the series of movies, the inheritance and changes in acting styles in different movies, the development context of the picture style, the penetration and innovation of the music style, etc. These elements are intertwined at the semantic level to construct a complex semantic network, and finally form the first template article semantic space. This space covers all possible semantic categories involved in the movie review template article when evaluating this movie series, from microscopic plot details to macroscopic cultural influence of the series of movies.
[0135] For the second template newspaper article (retrospective entertainment news template article about the acting career of the lead actor), the server also performs comprehensive semantic extraction. It will analyze semantic elements such as the representative works, role transitions, image shaping processes, status changes in the entertainment circle, and evolution of cooperation relationships with other actors and directors in each acting stage of the lead actor. These elements construct the second template article semantic space, which comprehensively presents the semantic information architecture of the entertainment news template article when reporting the acting career of the lead actor.
[0136] Next, according to the deep learning network, perform semantic matching analysis on the first template article semantic space to generate weight coefficients of each first semantic vector unit in the first template article semantic space, and perform semantic matching analysis on the second template article semantic space to generate weight coefficients of each second semantic vector unit in the second template article semantic space.
[0137] In deep learning networks, a large number of semantic models and algorithms related to movie reviews and entertainment news have been pre-trained. For each first semantic vector unit in the semantic space of the first template article, for example, in a retrospective movie review template article about a movie series, for the semantic vector unit regarding the plot features of the movie, the deep learning network will analyze the degree of closeness of its association with the overall semantics of the movie review. If the plot feature is one of the core attractions of the movie series and is in a key elaboration part in the entire movie review template article, then the weight coefficient of this semantic vector unit will be relatively high. This means that it has a greater impact on the semantic association of the template newspaper article. For example, if the plot twist in a certain movie is a classic element of the series and is repeatedly emphasized in the movie review and linked to the success of the movie, then the semantic vector unit related to this plot twist will be assigned a relatively high weight coefficient.
[0138] The same applies to each second semantic vector unit in the semantic space of the second template article. In a retrospective entertainment news template article about the acting career of a lead actor, for the semantic vector unit related to the representative works of the lead actor in a certain important acting stage, if this work is of milestone significance to the acting career of the lead actor and is the key object of description in the entertainment news template article, then its weight coefficient will be relatively high. For example, if the lead actor won an international award for a certain movie, the influence coefficient of the semantic vector unit related to this movie in the semantic association of the entire entertainment news template article will be very large, and thus it will be assigned a relatively high weight coefficient.
[0139] Then, based on the weight coefficients of each first semantic vector unit, semantic space optimization is performed on the semantic space of the first template article to generate x first semantic vector units, and based on the weight coefficients of each second semantic vector unit, semantic space optimization is performed on the semantic space of the second template article to generate x second semantic vector units.
[0140] Suppose that through calculation or according to preset rules, the value of x is determined (the determination method of x here is similar to the processing method of previous movie review articles and entertainment news articles and may be related to factors such as the semantic association degree threshold and the dimension of the semantic space of the template article). For the semantic space of the first template article, the first x first semantic vector units are selected in descending order of the weight coefficients of each first semantic vector unit. For example, in a retrospective movie review template article about a movie series, the first x semantic vector units with the highest weight coefficients may be the semantic vector units related to the plots of the most core works of the movie series, the most representative performances, the most unique picture styles, and the elements that have a decisive impact on the theme of the movie series.
[0141] For the semantic space of the second template article, the top x second semantic vector units are also selected according to the descending order of the weight coefficients of each second semantic vector unit. For example, in the template article of the retrospective entertainment news about the leading actor's acting career, they may be the semantic vector units related to the leading actor's several most famous representative works, the roles that have a significant impact on his image shaping, and key cooperation relationships, etc.
[0142] After that, according to the semantic matching analysis network, semantic matching analysis is performed on the x first semantic vector units and the x second semantic vector units to generate the first feature identifiers corresponding to the x first semantic vector units and the second feature identifiers corresponding to the x second semantic vector units.
[0143] The semantic matching analysis network will perform detailed semantic comparison and correlation analysis on the x first semantic vector units and the x second semantic vector units. Take the x first semantic vector units in the template article of the retrospective film review of a film series and the x second semantic vector units in the template article of the retrospective entertainment news about the leading actor's acting career as an example. For a semantic vector unit in the x first semantic vector units related to a classic plot of a certain film in the film series, the semantic matching analysis network will search for the part with semantic relevance to it among the x second semantic vector units. If in the template article of the retrospective entertainment news about the leading actor's acting career, the leading actor's acting experience in this film is an important turning point in his acting career, then in the semantic matching analysis process, this second semantic vector unit related to the leading actor's acting experience has semantic relevance to the first semantic vector unit related to the film plot.
[0144] Through such semantic matching analysis, a first feature identifier is generated for the x first semantic vector units, and this first feature identifier can locate the y first semantic vector units among the x first semantic vector units that have semantic relevance to the x second semantic vector units. For example, the first feature identifier may mark which semantic vector units in terms of plot, performance, picture style, etc. in the template article of the retrospective film review of a film series have semantic connections with the relevant semantic vector units in the template article of the retrospective entertainment news about the leading actor's acting career. Similarly, a second feature identifier is generated for the x second semantic vector units to locate the z second semantic vector units among the x second semantic vector units that have semantic relevance to the x first semantic vector units.
[0145] Next, the deep learning network is optimized according to the weight coefficients of each first semantic vector unit and the weight coefficients of each second semantic vector unit, and the semantic matching analysis network is optimized according to the first feature identifier and the second feature identifier.
[0146] In terms of optimizing the deep learning network based on the weight coefficients of each first semantic vector unit and each second semantic vector unit. First, determine the first training cost based on the error between the feature identifier reflecting whether there is a theme feature vector in the first semantic vector unit in the first template article semantic space and the weight coefficient of the first semantic vector unit in the first template article semantic space. For example, in a template article of a retrospective movie review of a movie series, if there is a first semantic vector unit for a deep exploration of the theme of the movie series, its weight coefficient represents its importance in the semantic association of the movie review. The feature identifier reflecting whether there is a theme feature vector may be judged based on some predefined rules or models. If there is a large error between the weight coefficient of this semantic vector unit and the expected weight coefficient judged according to its theme feature vector, for example, the actual weight coefficient is lower than the expected weight coefficient, this means that there may be a deviation in the deep learning network's judgment of the importance of this semantic vector unit, thus generating a certain training cost.
[0147] Similarly, determine the second training cost based on the error between the feature identifier reflecting whether there is a theme feature vector in the second semantic vector unit in the second template article semantic space and the weight coefficient of the second semantic vector unit in the second template article semantic space. For example, in a template article of a retrospective entertainment news of a leading actor's acting career, for a second semantic vector unit related to the core elements of the leading actor's image building, if its weight coefficient is inconsistent with the expected weight coefficient judged according to its theme feature vector, a second training cost will be generated.
[0148] Then, optimize the deep learning network based on the first training cost and the second training cost. This optimization process may involve adjusting various parameters in the deep learning network, such as the weights and biases of the neural network. If the first training cost is large, it indicates that there are more problems in the judgment of the weight coefficients of the semantic vector units in the first template article semantic space, and the network will adjust the parameters in the direction of reducing this training cost. For example, if the large first training cost is caused by inaccurate judgment of the weight coefficients of the semantic vector units related to the movie plot features, the network may adjust the parameters related to the semantic analysis of plot features to more accurately determine the weight coefficients.
[0149] In terms of optimizing the semantic matching analysis network based on the first feature identifier and the second feature identifier, the third training cost is determined based on the error between the first feature identifier and the semantically related feature identifier, and the fourth training cost is determined based on the error between the second feature identifier and the semantically related feature identifier. The semantically related feature identifier is a feature identifier that reflects whether there is a topic feature vector in the semantically relevant part of x first semantic vector units and x second semantic vector units. For example, if there is a deviation between the semantic relevance of the semantic vector units marked by the first feature identifier and related to the movie plot and the lead actor's performance and the expected semantic relevance according to the semantically related feature identifier, the third training cost will be generated.
[0150] Similarly, for the second feature identifier, if the semantic relevance of the semantic vector units marked as the lead actor's representative works and related to the movie picture style is inconsistent with the expectation of the semantically related feature identifier, the fourth training cost will be generated. Finally, the semantic matching analysis network is optimized based on the third training cost and the fourth training cost. This optimization process may involve adjusting parameters such as the matching algorithm and the semantic distance calculation method in the semantic matching analysis network to improve the accuracy of semantic matching analysis, so that the first feature identifier and the second feature identifier can more accurately reflect the semantic relevance between x first semantic vector units and x second semantic vector units. By continuously optimizing the deep learning network and the semantic matching analysis network in this way, the accuracy and efficiency of semantic analysis of template newspaper articles can be improved, and thus a better model and algorithm basis can be provided for processing the semantic analysis of actual newspaper articles.
[0151] Figure 2 The hardware structure diagram of the abstract generation system 100 provided by the embodiment of the present application for implementing the above-mentioned abstract generation method applied to newspaper article semantic mining is shown, as Figure 2 shown, the abstract generation system 100 may include a processor 110, a machine-readable storage medium 120, a bus 130, and a communication unit 140.
[0152] In one possible design, the abstract generation system 100 can be a single server or a server group. The server group can be centralized or distributed (e.g., the abstract generation system 100 can be a distributed system). In some embodiments, the abstract generation system 100 can be local or remote. For example, the abstract generation system 100 can access information and / or data stored in the machine-readable storage medium 120 via a network. Alternatively, the abstract generation system 100 can be directly connected to the machine-readable storage medium 120 to access the stored information and / or data. In some embodiments, the abstract generation system 100 can be implemented on the abstract generation system. By way of example only, the abstract generation system can include a private cloud, a semantic-related cloud, a hybrid cloud, a community cloud, a distributed cloud, an internal cloud, a multi-layer cloud, etc., or any aggregation thereof.
[0153] The machine-readable storage medium 120 can store data and / or instructions. In some embodiments, the machine-readable storage medium 120 can store data obtained from an external terminal. In some embodiments, the machine-readable storage medium 120 can store the data and / or instructions that the abstract generation system 100 uses to execute or complete the exemplary methods described in this application.
[0154] In a specific implementation process, one or more processors 110 execute the computer-executable instructions stored in the machine-readable storage medium 120, so that the processors 110 can execute the abstract generation method applied to the semantic mining of newspaper articles as described in the above method embodiments. The processors 110, the machine-readable storage medium 120, and the communication unit 140 are connected via a bus 130. The processors 110 can be used to control the sending and receiving operations of the communication unit 140.
[0155] For the specific implementation process of the processors 110, reference can be made to the various method embodiments executed by the above abstract generation system 100. Their implementation principles and technical effects are similar, and will not be elaborated here in this embodiment.
[0156] In addition, an embodiment of the present application also provides a readable storage medium, in which computer-executable instructions are set. When a processor executes the computer-executable instructions, the abstract generation method applied to the semantic mining of newspaper articles as described above is implemented.
[0157] It should be noted that, in order to simplify the presentation of the disclosure of the present application and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present application, sometimes multiple features are merged into one embodiment, drawing, or description thereof. Similarly, it should be noted that, in order to simplify the presentation of the disclosure of the present application and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present application, sometimes multiple features are merged into one embodiment, drawing, or description thereof.
Claims
1. A summary generation method applied to semantic mining of newspaper articles, characterized in that: The method comprises: Performing semantic extraction on the first newspaper article to generate a first article semantic space, and performing semantic extraction on the second newspaper article to generate a second article semantic space; Performing semantic matching analysis on the first article semantic space to generate weight coefficients of each first semantic vector unit in the first article semantic space, and performing semantic matching analysis on the second article semantic space to generate weight coefficients of each second semantic vector unit in the second article semantic space, wherein the weight coefficients reflect the influence coefficients on the semantic association of newspaper articles; According to the weight coefficients of the first semantic vector units, the first article semantic space is optimized to generate x first semantic vector units, and according to the weight coefficients of the second semantic vector units, the second article semantic space is optimized to generate x second semantic vector units; Performing semantic matching analysis on the x first semantic vector units and the x second semantic vector units to generate a first semantic association summary result between the first newspaper article and the second newspaper article; The step of performing semantic matching analysis on the x first semantic vector units and the x second semantic vector units to generate a first semantic association summary result between the first newspaper article and the second newspaper article includes: Performing semantic matching analysis on the x first semantic vector units and the x second semantic vector units to generate first feature identifiers corresponding to the x first semantic vector units and second feature identifiers corresponding to the x second semantic vector units, wherein the first feature identifier is used to locate y first semantic vector units among the x first semantic vector units that have semantic relevance to the x second semantic vector units, and the second feature identifier is used to locate z second semantic vector units among the x second semantic vector units that have semantic relevance to the x first semantic vector units, wherein the first feature identifier and the second feature identifier are used to represent corresponding masks; A semantic association analysis is performed on the y first semantic vector units and the z second semantic vector units to generate a first semantic association summary result between the first newspaper article and the second newspaper article.
2. The method for generating abstracts for semantic mining of newspaper articles according to claim 1, characterized in that: The method of performing semantic space optimization on the first article semantic space according to the weight coefficients of the first semantic vector units to generate x first semantic vector units, and performing semantic space optimization on the second article semantic space according to the weight coefficients of the second semantic vector units to generate x second semantic vector units includes: Determine x according to a semantic relevance threshold and a dimension of the first article semantic space, or determine x according to the semantic relevance threshold and a dimension of the second article semantic space; According to the weight coefficients of the first semantic vector units, based on the order of decreasing the weight coefficients, the first x first semantic vector units in the semantic space of the first article are output as the x first semantic vector units; and according to the weight coefficients of the second semantic vector units, based on the order of decreasing the weight coefficients, the first x second semantic vector units in the semantic space of the second article are output as the x second semantic vector units.
3. The method for generating abstracts for semantic mining of newspaper articles according to any one of claims 1 to 2, characterized in that: The performing semantic matching analysis on the x first semantic vector units and the x second semantic vector units to generate first feature identifiers corresponding to the x first semantic vector units and second feature identifiers corresponding to the x second semantic vector units includes: Determine the first basic feature identifiers corresponding to the x first semantic vector units, and determine the second basic feature identifiers corresponding to the x second semantic vector units; Filtering the x first semantic vector units according to the first basic feature identifier to generate x first semantic vector units to be optimized, and filtering the x second semantic vector units according to the second basic feature identifier to generate x second semantic vector units to be optimized; Performing semantic relevance analysis on the x first semantic vector units to be optimized according to the first semantic relevance analysis module, and generating x first adjusted semantic vector units based on the x second semantic vector units to be optimized and the x semantic vector units generated by the first semantic relevance analysis module according to the first interactive relevance analysis module, and performing semantic relevance analysis on the x second semantic vector units to be optimized according to the second semantic relevance analysis module, and generating x second adjusted semantic vector units based on the x first semantic vector units to be optimized and the x semantic vector units generated by the second semantic relevance analysis module; According to the x first adjustment semantic vector units, adjusting the first basic feature identifier to generate a first adjustment feature identifier, and according to the x second adjustment semantic vector units, adjusting the second basic feature identifier to generate a second adjustment feature identifier; The first feature identifier is determined according to the first adjustment feature identifier, and the second feature identifier is determined according to the second adjustment feature identifier.
4. The method for generating abstracts for semantic mining of newspaper articles according to claim 3, characterized in that: Before performing semantic extraction on the first newspaper article to generate the first article semantic space, and performing semantic extraction on the second newspaper article to generate the second article semantic space, the method further includes: When the dimension of the first newspaper article and the dimension of the second newspaper article are inconsistent, converting the first newspaper article and the second newspaper article to the same dimension; The determining of the first basic feature identifiers corresponding to the x first semantic vector units, and the determining of the second basic feature identifiers corresponding to the x second semantic vector units, include: Determine the first basic feature identifier according to the vocabulary supplemented in the dimensional conversion of the first newspaper article, wherein the mark corresponding to the semantic vector of the vocabulary supplemented in the x first semantic vector units in the first basic feature identifier is non-core content; The second basic feature identifier is determined based on the vocabulary supplemented in the dimensional conversion of the second newspaper article, and the mark corresponding to the semantic vector of the vocabulary supplemented in the x second semantic vector units in the second basic feature identifier is non-core content.
5. The method for generating abstracts for semantic mining of newspaper articles according to claim 3, characterized in that: The adjusting the first basic feature identifier according to the x first adjustment semantic vector units to generate a first adjustment feature identifier, and adjusting the second basic feature identifier according to the x second adjustment semantic vector units to generate a second adjustment feature identifier, include: Performing semantic matching analysis on the x first adjusted semantic vector units to generate a confidence combination corresponding to each of the x first adjusted semantic vector units, and performing semantic matching analysis on the x second adjusted semantic vector units to generate a confidence combination corresponding to each of the x second adjusted semantic vector units, wherein the confidence combination reflects the confidence of the semantically related content and the confidence of the semantically non-related content; Determine x first semantic identification codes corresponding to the x first adjusted semantic vector units based on the confidence combinations corresponding to the first adjusted semantic vector units, and determine x second semantic identification codes corresponding to the x second adjusted semantic vector units based on the confidence combinations corresponding to the second adjusted semantic vector units, wherein the first semantic identification code reflects the semantic content label corresponding to the maximum confidence in the confidence combinations corresponding to the first adjusted semantic vector units, and the second semantic identification code reflects the semantic content label corresponding to the maximum confidence in the confidence combinations corresponding to the second adjusted semantic vector units; Performing logic operation processing on the x first semantic identification codes to generate x first confidences, and performing logic operation processing on the x second semantic identification codes to generate x second confidences; The x first confidences are screened according to the first basic feature identifier to generate the first adjusted feature identifier, and the x second confidences are screened according to the second basic feature identifier to generate the second adjusted feature identifier.
6. The method for generating abstracts for semantic mining of newspaper articles according to any one of claims 1 to 2, characterized in that: The step of performing semantic extraction on the first newspaper article to generate the first article semantic space, and performing semantic extraction on the second newspaper article to generate the second article semantic space, comprises: Performing semantic extraction on the first newspaper article to generate the first article semantic space and the third article semantic space, and performing semantic extraction on the second newspaper article to generate the second article semantic space and the fourth article semantic space, wherein the dimension of the first article semantic space is different from the dimension of the third article semantic space, and the dimension of the second article semantic space is different from the dimension of the fourth article semantic space; After generating a first semantic association summary result between the first newspaper article and the second newspaper article, the method further includes: According to the first semantic association summary result, the dimension of the first article semantic space, the dimension of the third article semantic space and the third article semantic space, the third article semantic space is optimized to generate y newspaper article segments corresponding to the y first semantic vector units; and according to the first semantic association summary result, the dimension of the second article semantic space, the dimension of the fourth article semantic space and the fourth article semantic space, the fourth article semantic space is optimized to generate z newspaper article segments corresponding to the z second semantic vector units; According to the y newspaper article segments and the z newspaper article segments, the y newspaper article segments are semantically enhanced to generate y enhanced segments, and according to the y newspaper article segments and the z newspaper article segments, the z newspaper article segments are semantically enhanced to generate z enhanced segments; Extracting the core semantics of the y enhanced segments to generate y core semantics, and performing semantic compression or expansion processing on the z enhanced segments according to the semantic depth error between the third article semantic space and the fourth article semantic space to generate z adjusted segments; Performing semantic positioning processing on the z adjustment segments to generate z positioning semantics corresponding to the z adjustment segments; Performing semantic association analysis on the y core semantics and the z positioning semantics to generate a second semantic association summary result between the first newspaper article and the second newspaper article; The performing semantic positioning processing on the z adjustment segments to generate z positioning semantics corresponding to the z adjustment segments includes: Obtaining the semantic importance coefficient of each semantic unit in the first adjustment segment of the z adjustment segments; The semantic importance coefficient and the semantic content of each semantic unit in the first adjustment segment are weightedly fused to generate the positioning semantics of the first adjustment segment.
7. The method for generating abstracts for semantic mining of newspaper articles according to any one of claims 1 to 2, characterized in that: The method further comprises: Performing semantic extraction on the first template newspaper article to generate the semantic space of the first template article, and performing semantic extraction on the second template newspaper article to generate the semantic space of the second template article; According to the deep learning network, a semantic matching analysis is performed on the semantic space of the first template article to generate a weight coefficient of each first semantic vector unit in the semantic space of the first template article; and according to the deep learning network, a semantic matching analysis is performed on the semantic space of the second template article to generate a weight coefficient of each second semantic vector unit in the semantic space of the second template article, wherein the weight coefficient reflects an influence coefficient on the semantic association of the template newspaper article; According to the weight coefficients of the first semantic vector units, the semantic space of the first template article is optimized to generate x first semantic vector units, and according to the weight coefficients of the second semantic vector units, the semantic space of the second template article is optimized to generate x second semantic vector units; According to the semantic matching analysis network, a semantic matching analysis is performed on the x first semantic vector units and the x second semantic vector units to generate a first feature identifier corresponding to the x first semantic vector units and a second feature identifier corresponding to the x second semantic vector units, wherein the first feature identifier is used to locate y first semantic vector units having semantic relevance with the x second semantic vector units among the x first semantic vector units, and the second feature identifier is used to locate z second semantic vector units having semantic relevance with the x first semantic vector units among the x second semantic vector units; The deep learning network is optimized according to the weight coefficients of the first semantic vector units and the weight coefficients of the second semantic vector units, and the semantic matching analysis network is optimized according to the first feature identifier and the second feature identifier.
8. The method for generating abstracts for semantic mining of newspaper articles according to claim 7, characterized in that: The step of optimizing the deep learning network according to the weight coefficients of the first semantic vector units and the weight coefficients of the second semantic vector units includes: Determine a first training cost based on an error between a feature identifier reflecting whether a first semantic vector unit in a semantic space of a first template article has a topic feature vector and a weight coefficient of the first semantic vector unit in the semantic space of the first template article; and determine a second training cost based on an error between a feature identifier reflecting whether a second semantic vector unit in a semantic space of a second template article has a topic feature vector and a weight coefficient of the second semantic vector unit in the semantic space of the second template article; Optimizing the deep learning network according to the first training cost and the second training cost; The step of optimizing the semantic matching analysis network according to the first feature identifier and the second feature identifier includes: Determine a third training cost based on an error between the first feature identifier and the semantically relevant feature identifier, and determine a fourth training cost based on an error between the second feature identifier and the semantically relevant feature identifier; wherein the semantically relevant feature identifier is a feature identifier that reflects whether a topic feature vector exists in the semantically relevant part of the x first semantic vector units and the x second semantic vector units; The semantic matching analysis network is optimized according to the third training cost and the fourth training cost.
9. A summary generation system, characterized in that: The summary generation system includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the summary generation method applied to semantic mining of newspaper articles as described in any one of claims 1 to 8 above.
Citation Information
Patent Citations
Paper cold start disambiguation method based on feature extraction and fusion
CN115688737A
Natural language analysis system, and natural language analysis method
JP2014013549A