Sci-tech intelligence deep analysis method and system based on cross-modal semantic enhancement
By constructing a cross-modal semantic anchor set and semantic transmission path, bidirectional transmission and hierarchical parsing of textual and visual information are achieved, solving the problem of insufficient cross-modal information parsing in existing technologies and improving the quality and efficiency of scientific and technological intelligence parsing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-16
- Publication Date
- 2026-03-31
AI Technical Summary
Existing methods for analyzing scientific and technological intelligence fail to fully explore the semantic relationships between text and image information when processing cross-modal information. This makes it difficult to accurately grasp the intrinsic connections between different modalities and to effectively integrate cross-modal information to comprehensively reveal the core content, technological connections, and development trends of scientific and technological intelligence.
A cross-modal semantic anchor set is constructed, and bidirectional information transmission between text topic anchors and visual object anchors is realized through semantic transmission paths. A cross-modal semantic enhancement representation containing bidirectional enhancement information is generated, and hierarchical semantic parsing is performed. The topic association rules, technical element dependencies and concept evolution sequences are integrated to generate a structured science and technology intelligence analysis report.
It enhances the richness and accuracy of information expression, delves deeper into the inherent laws of scientific and technological intelligence, forms comprehensive analytical conclusions, optimizes cross-modal semantic association, and improves the quality and efficiency of scientific and technological intelligence analysis.
Smart Images

Figure CN121303139B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of scientific and technological intelligence analysis technology, and more specifically, to a method and system for deep analysis of scientific and technological intelligence based on cross-modal semantic enhancement. Background Technology
[0002] In the field of science and technology intelligence analysis, with the diversified development of information technology, science and technology intelligence exhibits cross-modal characteristics, encompassing various forms such as text and images. Existing science and technology intelligence analysis methods have significant limitations in processing cross-modal information.
[0003] On the one hand, traditional methods often process text and image information in isolation, failing to fully explore the semantic connections between them. For example, when analyzing scientific and technological literature, only the textual content is considered, while the key information contained in the illustrations is ignored, resulting in an incomplete understanding of scientific and technological intelligence. On the other hand, for the fusion of cross-modal information, existing technologies mostly adopt simple splicing or overlay methods, failing to achieve deep semantic interaction between text and image information. This makes it difficult to accurately grasp the intrinsic connections between different modalities when analyzing scientific and technological intelligence, and it is impossible to effectively integrate cross-modal information to comprehensively reveal the core content, technological connections, and development trends of scientific and technological intelligence, thus failing to meet the needs for in-depth analysis of scientific and technological intelligence in today's complex scientific and technological environment. Summary of the Invention
[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a method for deep analysis of scientific and technological intelligence based on cross-modal semantic enhancement, the method comprising:
[0005] Construct a cross-modal semantic anchor set, which includes text theme anchors extracted from scientific and technological intelligence texts, visual object anchors extracted from scientific and technological intelligence images, and the association mapping relationship between text theme anchors and visual object anchors. The text theme anchors carry semantic description information representing the core theme of scientific and technological intelligence, and the visual object anchors carry visual attribute information representing the key objects of scientific and technological intelligence.
[0006] Based on a set of cross-modal semantic anchors, a semantic transmission path between anchors is constructed. Through the semantic transmission path, bidirectional information transmission between text topic anchors and visual object anchors is realized. The semantic description information of text topic anchors is injected into visual object anchors, and the visual attribute information of visual object anchors is injected into text topic anchors, generating a cross-modal semantically enhanced representation containing bidirectional enhancement information.
[0007] A hierarchical semantic parsing method is used to perform cross-modal semantic enhancement representation. At the topic level, the association patterns between different text topic anchors are identified to form topic association rules. At the technology level, the technology element dependency mode corresponding to the technology-related anchors is analyzed to form technology element dependency relationship. At the evolution level, the performance changes of the same anchor in different scientific and technological intelligence fragments are tracked to form a concept evolution sequence. The topic association rules, technology element dependency relationship and concept evolution sequence are integrated to obtain the scientific and technological intelligence parsing conclusion.
[0008] The topic association rules in the analysis of scientific and technological intelligence are back-mapped to the cross-modal semantic anchor set. The strength of the association mapping relationship between the corresponding text topic anchor and the visual object anchor is adjusted according to the number of association basis clauses, the proportion of paragraphs where the association occurs, and the number of sub-topics covered by the association. Through multiple rounds of strength calibration, associations that exceed the reasonable association range are eliminated, and the updated cross-modal semantic anchor set is obtained.
[0009] Based on the updated cross-modal semantic anchor set, topic association rules, technical element dependencies, and concept evolution sequences are mapped to different module positions in the cross-modal semantic anchor set. Logical connection between modules is achieved through the association mapping relationship between anchors, and explanatory statements on the association between modules are added to generate a structured science and technology intelligence analysis report.
[0010] Furthermore, embodiments of the present invention also provide a deep analysis system for scientific and technological intelligence based on cross-modal semantic enhancement, characterized in that it includes:
[0011] A processor; a machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to execute the aforementioned deep analysis method for scientific and technological intelligence based on cross-modal semantic enhancement by executing the machine-executable instructions.
[0012] Based on the above, by constructing a cross-modal semantic anchor set, thematic anchors in scientific and technological intelligence texts and visual object anchors in images are accurately extracted, and the correlation mapping relationship between the two is clarified. Based on this cross-modal semantic anchor set, a semantic transmission path between anchors is constructed to realize bidirectional information transmission between text and visual information, generating a cross-modal semantic enhancement representation containing bidirectional enhancement information. This effectively integrates the semantic features of different modalities, improving the richness and accuracy of information expression. Layered semantic analysis is performed on the cross-modal semantic enhancement representation, deeply exploring the inherent laws of scientific and technological intelligence from the three levels of theme, technology, and evolution, forming a comprehensive analysis conclusion. By back-mapping the analysis conclusion to the cross-modal semantic anchor set and adjusting the strength of the correlation mapping relationship, the cross-modal semantic association is further optimized. Finally, a structured scientific and technological intelligence analysis report is generated based on the updated set, realizing the logical connection of each module, providing users with clear, accurate, and comprehensive scientific and technological intelligence analysis results, and significantly improving the quality and efficiency of scientific and technological intelligence analysis. Attached Figure Description
[0013] Figure 1 This is a schematic diagram of the execution flow of the deep analysis method for scientific and technological intelligence based on cross-modal semantic enhancement provided in an embodiment of the present invention.
[0014] Figure 2 This is a schematic diagram of exemplary hardware and software components of the deep analysis system for scientific and technological intelligence based on cross-modal semantic enhancement provided in an embodiment of the present invention. Detailed Implementation
[0015] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a flowchart illustrating a method for deep analysis of scientific and technological intelligence based on cross-modal semantic enhancement, provided in one embodiment of the present invention. The following is a detailed description of this method for deep analysis of scientific and technological intelligence based on cross-modal semantic enhancement.
[0016] Step S110: Construct a cross-modal semantic anchor set, which includes text theme anchors extracted from scientific and technological intelligence texts, visual object anchors extracted from scientific and technological intelligence images, and the association mapping relationship between text theme anchors and visual object anchors. The text theme anchors carry semantic description information representing the core theme of scientific and technological intelligence, and the visual object anchors carry visual attribute information representing the key objects of scientific and technological intelligence.
[0017] In this embodiment, the scenario of analyzing scientific and technological intelligence related to "deep learning model optimization technology" in the field of artificial intelligence is used. This intelligence includes the text content of three relevant academic papers and five experimental data charts. The process of constructing the cross-modal semantic anchor set described above requires extracting key anchors from both text and image modalities and establishing the correlation mapping relationship between them, thereby forming a structured set that can comprehensively reflect the core content of the scientific and technological intelligence.
[0018] First, the scientific and technological intelligence text is processed by extracting textual theme anchors through a series of text analysis techniques. These anchors must accurately represent the core theme of the scientific and technological intelligence, and their semantic descriptions should include the anchor's core meaning and category. Next, the scientific and technological intelligence images are processed by extracting visual object anchors using image processing techniques. These anchors should represent key objects in the image, and their visual attributes should include features such as shape, color, and texture. Finally, through specific association rules and verification mechanisms, the mapping relationship between the textual theme anchors and the visual object anchors is determined, ensuring that both semantically point to consistent scientific and technological intelligence content.
[0019] Step S111: Read the complete content of the science and technology intelligence text, and divide the science and technology intelligence text into multiple text paragraph units according to the semantic logical relationship between paragraphs. Each text paragraph unit corresponds to a sub-topic of the science and technology intelligence.
[0020] When reading scientific and technological information texts, it is necessary to obtain all character data of the text, including the title, abstract, body text, references, and other parts, to ensure that no content that may contain core information is missed. The obtained text content needs to undergo preliminary formatting to remove irrelevant formatting marks, such as headers, footers, and special symbols, to standardize the text data.
[0021] When splitting text according to the semantic and logical relationships between paragraphs, the first step is to identify topic sentences and transitional sentences. Topic sentences typically summarize the core content of a paragraph, while transitional sentences connect different topics. By analyzing these sentences, it can be determined whether paragraphs revolve around the same sub-topic. When there is a significant semantic shift between paragraphs, such as from introducing background knowledge to explaining specific methods, they can be divided into different text paragraph units. Each text paragraph unit should have a clear sub-topic, such as "Research Background of Deep Learning Model Optimization," "Principles of Gradient Descent-Based Optimization Algorithms," or "Experimental Data Acquisition and Processing Methods," to facilitate in-depth analysis of each sub-topic later.
[0022] Step S112: Perform sentence splitting on each text paragraph unit, splitting the text paragraph unit into multiple independent text sentences, and removing duplicate sentences containing redundant information and sentences with complete semantics.
[0023] When splitting text paragraphs into sentences, punctuation marks are the primary basis. Normally, periods, question marks, and exclamation marks indicate the end of a sentence. However, in scientific and technological information texts, there may be special cases, such as sentences quoting literature or statements containing formulas. In these cases, it is necessary to consider the context to ensure that the split sentences retain their independence and integrity.
[0024] After the segmentation is complete, the statements need to be filtered. For duplicate statements containing redundant information, a text similarity comparison method is used to calculate the similarity between statements. When the similarity between two statements exceeds a preset threshold, they are identified as duplicate statements, and only one of them needs to be retained to avoid information redundancy. For semantically incomplete statements, syntactic structure analysis is used to check whether the statement contains a complete subject-verb-object structure or other necessary components. If a statement lacks key components and cannot express a complete meaning, it is removed to ensure that the retained statements accurately convey semantic information.
[0025] Step S113: Perform word segmentation on each retained text statement to obtain multiple word units. Filter out function word units without actual semantic meaning and low-frequency irrelevant word units that appear only once in the scientific and technological information text, and retain core word units with scientific and technological information semantic value.
[0026] When performing word segmentation, use word segmentation tools suitable for scientific and technical texts. These tools can identify technical terms and compound words in the field of science and technology, accurately splitting text sentences into multiple word units. During the segmentation process, the characteristics of scientific and technical words need to be considered. For example, some compound technical terms may consist of multiple words and should be split as a whole word unit.
[0027] When filtering out semantically meaningless function words, a pre-defined function word dictionary is used. This dictionary includes prepositions, conjunctions, and auxiliary words in Chinese. By matching the segmented vocabulary units with this dictionary, function words are filtered out. For low-frequency, irrelevant vocabulary units that appear only once in the scientific and technological information text, the frequency of each unit within the entire text is first calculated. Then, combined with a professional dictionary in the scientific and technological field, it is determined whether the low-frequency words are relevant to the topic of the scientific and technological information. If the low-frequency words are irrelevant to the topic, such as occasional colloquial expressions or descriptive words unrelated to technical content, they are filtered out. The retained core vocabulary units should be terms and concepts that reflect the core content of the scientific and technological information, such as "deep learning," "neural network," "optimization algorithm," and "convergence speed."
[0028] Step S114: Combine the core vocabulary units according to semantic relevance to form multiple vocabulary combinations that can represent the core meaning of the sub-topic, while retaining indivisible and semantically long sentence fragments. The vocabulary combinations and long sentence fragments constitute a set of candidate text topic anchors.
[0029] When grouping core vocabulary units according to semantic relevance, the semantic similarity between them must first be calculated. A word vector model can be used to convert vocabulary units into vector representations, and then the degree of semantic relevance can be measured by calculating the cosine similarity between the vectors. When two or more core vocabulary units have high semantic similarity and appear together in multiple sentences, they are grouped into a single vocabulary combination.
[0030] For example, the core lexical units "gradient descent" and "learning rate" appear simultaneously in multiple sentences and are semantically closely related, jointly describing the key parameters of the optimization algorithm. Therefore, they can be combined into the lexical combination "gradient descent learning rate". For some long sentence fragments that cannot be broken down into multiple independent core lexical units but are semantically complete, such as "image feature extraction method based on convolutional neural network", since it can already completely represent a specific technical concept, it can be directly retained as a candidate text topic anchor.
[0031] The candidate text topic anchor set should include all word combinations and long sentence fragments obtained through the above method. These candidate anchors will serve as the basis for subsequent screening of text topic anchors.
[0032] Step S115: For each candidate anchor in the candidate text theme anchor set, count the number of times each candidate anchor appears in all paragraphs of the science and technology intelligence text, count the number of times each candidate anchor appears in key paragraphs that describe core technologies, core viewpoints, and core conclusions, and count the number of keywords that overlap between each candidate anchor and the core theme description statement of the science and technology intelligence text. Sort the candidate anchors by priority in the order of the number of times they appear, the number of times they appear in key paragraphs, and the number of keywords that overlap.
[0033] Step S1151: Traverse all text paragraphs of the science and technology information text, count the number of times each candidate anchor appears in each text paragraph, and sum the occurrence counts in all text paragraphs to obtain the total occurrence count of each candidate anchor.
[0034] When traversing all text paragraphs of a scientific and technological intelligence text, it is necessary to check each paragraph unit to see if the sentences contain candidate anchors. For each candidate anchor, perform exact matching or fuzzy matching in each paragraph unit and count its occurrences. Exact matching requires that the word combination or long sentence fragment of the candidate anchor is completely consistent with the content in the paragraph; fuzzy matching allows for some parts of speech changes or word order adjustments, but the core vocabulary must be consistent.
[0035] The total number of occurrences of each candidate anchor point is calculated by summing the occurrence counts across all text paragraphs. This total count reflects the prevalence of the candidate anchor point within the entire scientific and technological information text; a higher occurrence count generally indicates greater importance of the candidate anchor point within the text.
[0036] Step S1152: Identify key paragraphs in the scientific and technological intelligence text. The key paragraphs include paragraphs that describe core technologies, paragraphs that describe core viewpoints, and paragraphs that describe core conclusions. Count the number of times each candidate anchor point appears in each key paragraph. Add up the number of times each candidate anchor point appears in all key paragraphs to obtain the total number of times each candidate anchor point appears in a key paragraph.
[0037] When identifying key paragraphs, first analyze the structure of the scientific and technological information text. Paragraphs describing the core technology are usually located in the methods section of the main text, detailing the proposed technical solutions, algorithm processes, etc. Paragraphs describing the core viewpoints may be distributed in the introduction, discussion, etc., expressing the author's views on the field and the innovative points of the research. Paragraphs describing the core conclusions are generally in the conclusion section, summarizing the main results and contributions of the research.
[0038] The locations of these key paragraphs are determined through manual annotation or keyword recognition. Then, the frequency of candidate anchor points in each key paragraph is counted, using either exact or fuzzy matching. The total frequency of occurrences in all key paragraphs is then summed to obtain the total number of occurrences. This total number of occurrences reflects the relevance of candidate anchor points to the core content of the scientific and technological intelligence; candidate anchor points that appear more frequently in key paragraphs are generally considered more important.
[0039] Step S1153: Extract the core theme description statements from the scientific and technological information text. The core theme description statements are located in the introduction, abstract, or conclusion of the text. The core theme description statements are used to summarize the overall core content of the scientific and technological information. Extract the keywords from the core theme description statements to form a list of core keywords.
[0040] When extracting core theme descriptions, focus on the introduction, abstract, and conclusion sections of the text. These sections typically provide a summary of the overall content of the scientific and technological information. The introduction may introduce the background, purpose, and significance of the research; the abstract highly condenses the main content, methods, results, and conclusions; and the conclusion summarizes the research findings and future prospects.
[0041] From these sections, statements that comprehensively summarize the core content of scientific and technological intelligence are selected as core theme description statements. Then, keywords are extracted from these core theme description statements using keyword extraction tools based on TF-IDF or TextRank algorithms to identify core words. These core words collectively constitute a core keyword list, such as "deep learning model," "optimization algorithm," "performance improvement," and "experimental verification."
[0042] Step S1154: Extract keywords from the semantic description of each candidate anchor point to form a candidate keyword list, and count the number of overlapping keywords between the candidate keyword list and the core keyword list.
[0043] For each candidate anchor point, its semantic description needs to be obtained first. For candidate anchor points in the form of word combinations, their semantic description can be inferred from the meaning of the combined words; for candidate anchor points in the form of long sentence fragments, their semantic description is the sentence itself. Then, using the same method as extracting the core keyword list, keywords are extracted from the semantic descriptions of the candidate anchor points to form a candidate keyword list.
[0044] Compare the candidate keyword list with the core keyword list and count the number of identical keywords. The more overlapping keywords, the closer the candidate anchor is to the core theme of the scientific and technological intelligence, and the higher its priority should be as a text theme anchor.
[0045] Step S1155: Using the total number of occurrences, the total number of occurrences in key paragraphs, and the number of overlapping keywords as ranking indicators, preliminary ranking of candidate anchors is performed according to the total number of occurrences from highest to lowest. If two candidate anchors have the same total number of occurrences, the total number of occurrences in key paragraphs of the two candidate anchors is compared, and the candidate anchor with more total occurrences in key paragraphs is ranked first. If the total number of occurrences and the total number of occurrences in key paragraphs of two candidate anchors are the same, the number of overlapping keywords of the two candidate anchors is compared, and the candidate anchor with more overlapping keywords is ranked first. If the total number of occurrences, the total number of occurrences in key paragraphs, and the number of overlapping keywords of two candidate anchors are all the same, the semantic description length of the candidate anchors is compared, and the candidate anchor with the shortest semantic description length is ranked first. A priority ranking table for candidate anchors is generated, which includes the candidate anchor's identifier, the total number of occurrences of the candidate anchor, the total number of occurrences in key paragraphs of the candidate anchor, the number of overlapping keywords of the candidate anchor, and the candidate anchor's ranking.
[0046] When performing initial sorting, the total frequency of occurrence is the primary indicator. Candidate anchors with a high total frequency are more common in the text and are likely important concepts or terms. When the total frequency is the same, the total frequency in key paragraphs becomes the next comparison indicator, because candidate anchors appearing more frequently in key paragraphs are more closely related to the core content.
[0047] If the total frequency of occurrence and the total frequency of occurrence in key paragraphs are the same, then compare the number of overlapping keywords. Candidate anchors with more overlapping keywords have a higher degree of match with the core theme. When all three indicators are the same, compare the semantic description length of the candidate anchors. Candidate anchors with shorter semantic descriptions are usually more concise and clear, and can more directly represent the core meaning of the sub-theme.
[0048] Based on the above sorting rules, all candidate anchor points are sorted, and a priority sorting table is generated.
[0049] Step S116: Select the candidate anchor points with higher priority as the final text topic anchor points, and generate corresponding semantic description statements for each text topic anchor point. The semantic description statements fully cover the core semantics and sub-topic categories of the text topic anchor point.
[0050] When selecting the final text topic anchors, the number needs to be determined based on the complexity of the scientific and technological intelligence and the analytical requirements. Generally, the top 20%-30% of candidate anchors in the priority ranking table should be selected as text topic anchors to ensure comprehensive coverage of the core content of the scientific and technological intelligence, while avoiding excessive anchors that could complicate subsequent processing.
[0051] When generating semantic descriptions for each text topic anchor, it is necessary to consider its contextual information within the scientific and technological information text. The semantic description should accurately summarize the core semantics of the text topic anchor, clearly defining the concepts, principles, methods, etc., it should also indicate the subtopic to which the text topic anchor belongs, such as "belongs to the subtopic of optimization algorithm principles" or "belongs to the subtopic of experimental results analysis," making the semantic description more complete and clear. For example, for the candidate anchor "gradient descent learning rate," its semantic description could be: "This text topic anchor belongs to the subtopic of optimization algorithm principles, and its core semantics are the learning rate parameter used in the gradient descent optimization algorithm to adjust the parameter update step size."
[0052] Step S117: Load the complete pixel data of the science and technology intelligence image, and use a Gaussian filtering algorithm to remove visual noise from the pixel data, retaining the visual content that reflects the key objects of the science and technology intelligence.
[0053] When loading images for scientific and technological information, it is necessary to read the raw pixel data of the image, including information such as the image width, height, and color channels. For image files of different formats, such as JPEG and PNG, appropriate image reading libraries need to be used to ensure that the pixel values of the image can be accurately obtained.
[0054] When using Gaussian filtering for visual noise removal, the first step is to determine the filter parameters, such as the filter size and standard deviation. The filter size is typically chosen as an odd number, such as 3x3 or 5x5, while the standard deviation is adjusted based on the severity of the image noise. A larger standard deviation results in a stronger filtering effect but may lead to image blurring; a smaller standard deviation results in a weaker filtering effect and may not effectively remove noise.
[0055] When applying the Gaussian filtering algorithm, the pixel values are smoothed by convolving the filter with each pixel of the image. During the convolution process, the center pixel of the filter has the largest weight, and the weight of the surrounding pixels gradually decreases. This removes noise while preserving as much edge and detail information as possible, ensuring that the visual content reflecting key scientific and technological information, such as experimental data curves, coordinate axes in charts, and legends, is not excessively blurred.
[0056] Step S118: The denoised science and technology information image is divided into regions using a region growing-based visual region segmentation algorithm. Based on the similarity of pixel gray values, the science and technology information image is segmented into multiple non-overlapping visual regions, and each visual region corresponds to a potential visual object.
[0057] The basic idea of visual region segmentation algorithms based on region growing is to start with seed pixels in the image and merge adjacent pixels with similar gray values into the same region, gradually growing to form a complete visual region. First, suitable seed pixels need to be selected. Seed pixels can be determined manually or through automatic detection, and typically representative pixels, such as the center pixel of the region, are chosen.
[0058] After determining the seed pixels, a threshold for grayscale similarity is set. For each seed pixel, the difference between the grayscale value of its surrounding pixels and the grayscale value of the seed pixel is checked to see if it is less than the threshold. If it is less than the threshold, the neighboring pixel is merged into the current region and used as a new seed pixel to continue growth. This process is repeated until no new pixels can be merged into the region.
[0059] Using the method described above, the denoised scientific and technological information image is segmented into multiple non-overlapping visual regions. Pixels within each visual region have similar grayscale values, representing a potential visual object, such as the curved regions in the table, the bar regions in the histogram, and the background regions of the image.
[0060] Step S119: Perform object recognition for each visual region. By comparing the pixel distribution features of the visual region with the preset key object feature library of scientific and technological intelligence, retain the visual regions whose pixel distribution features match any feature in the key object feature library, and filter out the visual regions containing key information objects.
[0061] The pre-defined key object feature library for scientific and technological intelligence contains the pixel distribution features of common key objects in scientific and technological intelligence images. For example, the pixel distribution features of experimental data curves are usually represented as a continuous set of pixels with a certain slope; the pixel distribution features of coordinate axes are represented as a set of horizontal or vertical straight line pixels; and the pixel distribution features of legends include different colored or shaped labels and corresponding text description pixel areas, etc.
[0062] When performing object recognition on each visual region, the pixel distribution features of that region are first extracted, including the grayscale distribution, spatial location distribution, and shape features of the pixels. Then, the extracted pixel distribution features are compared with features in a key object feature library. During the comparison, a matching method based on feature vector similarity is used to calculate the similarity between the pixel distribution feature vector of the visual region and each feature vector in the key object feature library. When the similarity exceeds a preset matching threshold, the visual region is determined to match the corresponding feature in the key object feature library, and the visual region is retained; otherwise, it is considered a non-critical information region and removed.
[0063] Step S1110: For the visual region containing the key information object, extract the coordinate data of the boundary contour of the visual region to obtain shape features, extract the frequency and amplitude data of pixel grayscale value changes in the visual region to obtain texture features, and extract the RGB color value distribution ratio data of each pixel in the visual region to obtain color features.
[0064] When extracting the coordinate data of the visual region's boundary contour, an edge detection algorithm, such as the Canny edge detection algorithm, is used to identify the boundary pixels of the visual region. Then, a contour tracking algorithm is used to record the coordinate data of the boundary pixels in a certain order (such as clockwise or counterclockwise). The above coordinate data constitutes the boundary contour of the visual region. By analyzing the boundary contour, the shape features of the visual region can be obtained, such as the perimeter, area, and concavity / convexity of the contour.
[0065] When extracting texture features, statistical analysis is performed on the pixel grayscale values within the visual region. The frequency of grayscale value changes is calculated, i.e., the number of times the grayscale value changes per unit area; the amplitude of grayscale value changes is calculated, i.e., the difference between the maximum and minimum grayscale values. This data reflects the roughness or fineness of the texture within the visual region. For example, texture features on experimental data curves may exhibit gradual changes in grayscale values, while texture features of the image background may exhibit random changes in grayscale values.
[0066] When extracting color features, for color images, it is necessary to statistically analyze the distribution ratio of the RGB color values of each pixel within the visual region across different intervals. For example, the color values of the red channel can be divided into multiple intervals, and the proportion of pixels in each interval to the total number of pixels can be calculated. A similar statistical analysis can be performed on the green and blue channels. The aforementioned distribution ratio data constitutes the color features of the visual region, which can be used to distinguish key objects of different colors, as shown in the illustration of different color markers.
[0067] Step S1111: Determine the visual regions from which shape features, texture features, and color features have been extracted as visual object anchor points, generate corresponding visual attribute feature descriptions for each visual object anchor point, and record the coordinate position information of each visual object anchor point in the science and technology intelligence image.
[0068] Visual regions from which shape, texture, and color features are simultaneously extracted are identified as anchor points for visual objects. This is because these features can describe the attributes of a visual object from different perspectives: shape features reflect the object's geometric form, texture features reflect the object's surface details, and color features reflect the object's color information. The combination of these three features can comprehensively represent a visual object.
[0069] When generating visual attribute feature descriptions for each visual object anchor point, the extracted shape features, texture features, and color features need to be quantified and described in text. For example, shape features can be described as "the boundary contour is approximately a straight line, and the length is about two-thirds of the image width"; texture features can be described as "the pixel grayscale values change at a low frequency and with a small amplitude, and the overall texture is relatively smooth"; color features can be described as "the red channel accounts for about 30% of the RGB color values, the green channel accounts for about 50%, and the blue channel accounts for about 20%".
[0070] When recording the coordinate position information of the visual object anchor point in the scientific and technological information image, the upper left corner of the image is usually taken as the origin, and the coordinates of the upper left and lower right corners of the boundary rectangle of the visual area are recorded, such as "the coordinate position is (x1, y1) to (x2, y2)", so that the visual object anchor point can be accurately located and referenced later.
[0071] Step S1112: Establish the mapping relationship between the semantic description of the text topic anchor and the visual attribute feature description of the visual object anchor. Calculate the association similarity by counting the number of keyword overlaps between the semantic description and the visual attribute feature description, counting the number of corresponding terms in the preset semantic-visual association dictionary, and counting the number of matching terms in the overall semantic comparison of the two. Pair the text topic anchors with association similarity exceeding the preset association standard with the visual object anchors.
[0072] Step S11121: Extract the core keywords from the semantic description statement of each text topic anchor point. The core keywords are nouns, verbs or adjectives used to represent the core meaning of the semantic description statement. Remove function words and modifying words from the semantic description statement to form a list of text keywords.
[0073] When extracting core keywords from semantic description statements of text themes, the statements are first segmented into multiple lexical units. Then, based on the part of speech, nouns, verbs, and adjectives are selected, as these typically carry the core meaning of the statement. Simultaneously, function words such as prepositions, conjunctions, and auxiliary words, as well as some modifying words like "very," "relatively," and "to a certain extent," are removed, as these contribute less to the core semantic meaning.
[0074] The selected core keywords are combined to form a list of text keywords. For example, for the semantic description statement "This text's topic anchor point belongs to the subtopic of optimization algorithm principles, and the core semantics is the learning rate parameter used to adjust the parameter update step size in the gradient descent optimization algorithm," the extracted core keywords could be "optimization algorithm principles," "gradient descent," "parameter update step size," and "learning rate parameter," thus forming a list of text keywords.
[0075] Step S11122: Extract key feature words from the visual attribute feature description of each visual object anchor point. Key feature words are words used to characterize visual attribute features. Key feature words include words corresponding to shape features, words corresponding to color features, and words corresponding to texture features, forming a list of visual keywords.
[0076] When extracting key feature words from the visual attribute feature description of visual object anchor points, extraction is performed separately for shape features, color features, and texture features. Shape features correspond to terms such as "straight line," "curve," "rectangle," "circle," and "boundary contour"; color features correspond to terms such as "red," "green," "blue," "grayscale value," and "color distribution"; and texture features correspond to terms such as "smooth," "rough," "grayscale change frequency," and "grayscale change amplitude."
[0077] Collect the aforementioned key feature words to form a list of visual keywords. For example, for the visual attribute feature description "the boundary contour is approximately a straight line, the pixel grayscale value changes at a low frequency and with a small amplitude, the overall surface is relatively smooth, and the red channel accounts for about 30% of the RGB color values", the extracted key feature words could be "straight line", "boundary contour", "grayscale value change frequency", "grayscale change amplitude", "smooth", "red channel", and "proportion", thus forming a list of visual keywords.
[0078] Step S11123: Establish a semantic-visual association dictionary, which records the correspondence between preset semantic keywords of scientific and technological information and visual feature words.
[0079] Building a semantic-visual association dictionary requires a large amount of scientific and technological intelligence data and domain knowledge. First, common semantic keywords and their corresponding visual feature words in the scientific and technological intelligence field are collected. Semantic keywords include technical terms, concept names, method names, etc., while visual feature words include words describing shape, color, texture, etc. Then, the correspondence between semantic keywords and visual feature words is determined through manual annotation or machine learning methods.
[0080] For example, the semantic keyword "experimental data curve" may correspond to visual feature words such as "curve", "continuous pixels", and "gradual grayscale change"; the semantic keyword "coordinate axis" may correspond to visual feature words such as "straight line", "horizontal", "vertical", and "scale mark". The above correspondences are recorded in the semantic-visual association dictionary.
[0081] Step S11124: Count the number of overlapping keywords between the text keyword list and the visual keyword list. The number of overlapping keywords is the total number of words that exist in both the text keyword list and the visual keyword list.
[0082] The words in the text keyword list and the visual keyword list are compared one by one to identify the identical words in both lists. These identical words are called overlapping keywords. The total number of overlapping keywords is counted to obtain the number of overlapping keywords. The number of overlapping keywords reflects the degree of direct correlation at the lexical level between the semantic description of the text's topic anchor and the visual attribute feature description of the visual object anchor.
[0083] Step S11125: Based on the semantic-visual association dictionary, count the number of terms in the semantic-visual association dictionary that correspond to the keywords in the text keyword list and the keywords in the visual keyword list. The number of terms that correspond to the keywords is the total number of terms in which the text keywords and visual keywords form a correspondence.
[0084] Iterate through each keyword in the text keyword list and search for a corresponding visual keyword in the semantic-visual association dictionary. Then, check if these visual keywords exist in the visual keyword list. If they exist, it means that the text keyword corresponds to a visual keyword in the visual keyword list, and this is recorded as a correspondence entry.
[0085] The number of all the aforementioned corresponding clauses is counted to obtain the total number of corresponding clauses. For example, the text keyword "gradient descent" corresponds to the visual keyword "curve" in the semantic-visual association dictionary, and since the list of visual keywords contains "curve," a corresponding clause is formed. The number of corresponding clauses reflects the degree of indirect semantic association between the text topic anchor and the visual object anchor.
[0086] Step S11126: Using a semantic similarity calculation algorithm in natural language processing, perform an overall semantic comparison between the semantic description of the text topic anchor and the visual attribute feature description of the visual object anchor, and count the number of semantic matching clauses between the two. The number of semantic matching clauses is the total number of clauses in which the semantic description and the visual attribute feature description have corresponding semantic relationships.
[0087] This study employs semantic similarity calculation algorithms from natural language processing, such as those based on pre-trained language models (e.g., BERT). First, the semantic descriptions of text topic anchors and the visual attribute feature descriptions of visual object anchors are input into the pre-trained language model to obtain their semantic vector representations. Then, the cosine similarity between the two semantic vectors is calculated, and the overall semantic association between them is determined based on the similarity value.
[0088] Based on a preset semantic matching threshold, when the cosine similarity exceeds this threshold, it is determined that there is a semantic match between the semantic description and the visual attribute feature description. The number of all semantic matching terms that meet the condition is counted to obtain the total number of semantic matching terms. The total number of semantic matching terms reflects the degree of semantic association between the text's topic anchor and the visual object anchor.
[0089] Step S11127: Combine the number of overlapping keywords, the number of corresponding clauses, and the number of semantically matched clauses according to a preset ratio to calculate the association similarity between the semantic description statement and the visual attribute feature description.
[0090] The preset proportions are set based on the domain characteristics of scientific and technological intelligence and the importance of cross-modal associations. For example, the weight of the number of overlapping keywords can be set to 40%, the weight of the number of corresponding clauses to 30%, and the weight of the number of semantically matched clauses to 30%. Then, the number of overlapping keywords, the number of corresponding clauses, and the number of semantically matched clauses are multiplied by their respective weights, and the products are added together to obtain the comprehensive score of association similarity.
[0091] For example, if the number of overlapping keywords is 2, the number of corresponding relational clauses is 3, the number of semantically matched clauses is 1, and the preset ratios are 40%, 30%, and 30%, then the association similarity is (2×40%) + (3×30%) + (1×30%) = 0.8 + 0.9 + 0.3 = 2.0 (this is just an example; in actual calculations, normalization processing is required based on the specific numerical range to ensure that the association similarity is between 0 and 1).
[0092] Step S11128: Set a preset association standard for association similarity, wherein the preset association standard is determined based on the domain characteristics and cross-modal association requirements of scientific and technological intelligence.
[0093] Setting predefined association criteria requires consideration of the domain characteristics of scientific and technological information. The degree of association between text and images may vary across different domains. For example, in the field of computer vision, the association between visual objects in images and textual descriptions may be more direct and closer; while in the field of theoretical physics, images may serve more as supplementary explanations of textual content, with a relatively weaker degree of association.
[0094] Simultaneously, cross-modal association requirements also need to be considered. If higher association accuracy is required, the preset association standard should be set higher; if it is necessary to retain as many potential associations as possible, the preset association standard can be set lower. Typically, the preset association standard is determined through multiple experiments and data analysis, generally ranging from 0.5 to 0.7.
[0095] Step S11129: Pair text topic anchors with a similarity exceeding the preset association standard with visual object anchors, record the identification information and association similarity of each pair of anchors, and form a preliminary anchor pairing set.
[0096] Iterate through all combinations of text topic anchors and visual object anchors, calculating the correlation similarity between them. When the correlation similarity exceeds a preset correlation standard, pair the text topic anchor with the visual object anchor. Simultaneously, record the unique identifier information of each pair of anchors, such as the text topic anchor ID and the visual object anchor ID, as well as their correlation similarity value.
[0097] The above pairing information is compiled to form a preliminary anchor pairing set. This preliminary anchor pairing set may contain pairings of multiple text topic anchors with the same visual object anchor, or pairings of multiple visual object anchors with the same text topic anchor.
[0098] Step S111210: Perform deduplication on the initial anchor pairing set. If a text topic anchor is paired with multiple visual object anchors, or a visual object anchor is paired with multiple text topic anchors, retain the pairing combination with the highest association similarity and remove the other pairing combinations, so that each anchor corresponds to only one associated anchor in the pairing set.
[0099] During deduplication, the text is first grouped by its topic anchor ID. For each text topic anchor ID, all corresponding visual object anchor pairs are examined, and the pair with the highest correlation similarity is retained, while the rest are discarded. Then, the text is grouped again by its visual object anchor ID. For each visual object anchor ID, all corresponding text topic anchor pairs are examined, and again, the pair with the highest correlation similarity is retained, while the rest are discarded.
[0100] Through the above deduplication process, it is ensured that each text topic anchor is associated with only one visual object anchor in the pairing set, and each visual object anchor is also associated with only one text topic anchor, forming a one-to-one anchor pairing relationship and avoiding complex many-to-many associations.
[0101] Step S1113: Perform cross-modal semantic consistency verification on the paired anchor point combinations. By comparing the semantic meaning of the text topic anchor points with the scientific and technological information content pointed to by the visual representation of the visual object anchor points, eliminate anchor point combinations whose semantic meaning and visual representation point to different scientific and technological information content.
[0102] Cross-modal semantic consistency verification requires in-depth analysis of the semantic meaning of text topic anchors and the visual representation of visual object anchors. The semantic meaning of text topic anchors is reflected in their semantic descriptions, clearly defining the scientific and technological information they represent; the visual representation of visual object anchors is reflected in their visual attribute features, also pointing to specific scientific and technological information content.
[0103] The scientific and technological information they point to is compared to determine if they are consistent. For example, if the semantic meaning of a text topic anchor points to "the convergence speed of optimization algorithm A," and the visual representation of a visual object anchor points to "the experimental data curve of optimization algorithm A," then the scientific and technological information they point to is consistent. However, if the visual representation of a visual object anchor points to "the experimental data curve of optimization algorithm B," then the content they point to is inconsistent. For anchor combinations pointing to different scientific and technological information, they are removed from the pairing set to ensure that the retained anchor combinations are semantically consistent.
[0104] Step S1114: Integrate the verified text topic anchors, verified visual object anchors, and the association mapping relationship between the verified text topic anchors and visual object anchors to form a cross-modal semantic anchor set. Each anchor in the cross-modal semantic anchor set carries unique identification information to distinguish different anchors.
[0105] When integrating validated text topic anchors, visual object anchors, and associated mappings, firstly, all validated text topic anchors and visual object anchors are summarized separately to ensure the completeness of information for each anchor, including identification information, semantic descriptions, or visual attribute feature descriptions. Then, the validated associated mappings are linked to the corresponding anchors, clarifying which visual object anchor each text topic anchor is associated with.
[0106] Each anchor is assigned a unique identifier, such as "T-XXX" for text topic anchors and "V-XXX" for visual object anchors, where XXX is a numerical code. This ensures accurate differentiation of different anchors within the cross-modal semantic anchor set. Finally, this information is organized into a structured data format, such as JSON or XML, to form the cross-modal semantic anchor set.
[0107] Step S120: Construct a semantic transmission path between anchors based on the cross-modal semantic anchor set. Achieve bidirectional information transmission between text topic anchors and visual object anchors through the semantic transmission path. Inject the semantic description information of the text topic anchors into the visual object anchors, and at the same time inject the visual attribute information of the visual object anchors into the text topic anchors, thereby generating a cross-modal semantic enhancement representation containing bidirectional enhancement information.
[0108] Constructing semantic transmission paths between anchors based on a cross-modal semantic anchor set requires using text topic anchors and visual object anchors as nodes, and determining the information transmission path based on the association mapping relationship between them. The semantic transmission path should enable bidirectional information transmission; that is, the semantic description information of the text topic anchor can be transmitted to the visual object anchor, and the visual attribute information of the visual object anchor can also be transmitted to the text topic anchor.
[0109] During information transmission, it is necessary to filter and integrate the transmitted information to ensure that the information injected into the recipient's anchor is core and crucial. Injecting semantic description information from text topic anchors into visual object anchors enriches their semantic connotation; conversely, injecting visual attribute information from visual object anchors into text topic anchors provides intuitive visual support. Through this bidirectional information transmission, each anchor contains enhanced information from associated anchors, thereby generating a cross-modal semantically enhanced representation.
[0110] Step S121: Extract the association mapping relationship between all text topic anchors and visual object anchors from the cross-modal semantic anchor set, count the number of visual object anchors corresponding to each text topic anchor, count the number of text topic anchors corresponding to each visual object anchor, and obtain the connection relationship between anchors.
[0111] When extracting association mappings from a cross-modal semantic anchor set, it is necessary to traverse all association records in the set, obtain the visual object anchor ID corresponding to each text topic anchor ID, and the text topic anchor ID corresponding to each visual object anchor ID. Due to deduplication, each text topic anchor corresponds to one visual object anchor, and each visual object anchor also corresponds to one text topic anchor; therefore, the connection relationship is one-to-one.
[0112] The number of visual object anchors corresponding to each text topic anchor and the number of text topic anchors corresponding to each visual object anchor should both be 1. This statistical analysis confirms the correctness of the connections between anchors and identifies any unrelated or incorrectly associated anchors. The connections between anchors form the foundation for constructing a semantic transmission path network.
[0113] Step S122: Using text topic anchors and visual object anchors as nodes and the association mapping relationship between anchors as edges, construct a semantic transmission path network between anchors. Each edge in the semantic transmission path network between anchors corresponds to a semantic transmission path from a text topic anchor to a visual object anchor or from a visual object anchor to a text topic anchor.
[0114] When constructing the semantic transmission path network between anchor points, text topic anchor points and visual object anchor points are abstracted as nodes in the network. Each node has unique identification information, corresponding to a specific anchor point. The association mapping relationship between anchor points is used as edges connecting the nodes, and each edge represents a bidirectional semantic transmission path between text topic anchor points and visual object anchor points.
[0115] In the network, the edges are bidirectional, meaning information can be transmitted along the edges from text topic anchors to visual object anchors, and vice versa. Thus, the entire network forms a graph structure composed of nodes and bidirectional edges.
[0116] Step S123: Analyze the semantic contribution of each text theme anchor in the scientific and technological information text, count the number of times the text theme anchor appears in the core paragraph, count the number of associations between the text theme anchor and other text theme anchors, count the number of clauses of the text theme anchor that explain the core meaning of the text, standardize the semantic contribution, number of occurrences, number of associations, and number of clauses respectively, and merge the standardized semantic contribution, number of occurrences, number of associations, and number of clauses according to a preset ratio to obtain the semantic importance assessment result of the text theme anchor.
[0117] When analyzing the semantic contribution of thematic anchors in a text, it is necessary to consider their importance to the expression of the core content within the entire scientific and technological intelligence text. Semantic contribution can be determined through expert evaluation or based on semantic analysis algorithms; the higher the semantic contribution, the more important the anchor is to understanding the core meaning of the text.
[0118] The analysis includes counting the frequency of text topic anchors in core paragraphs, such as "Method Proposal," "Experimental Results Analysis," and "Conclusion." A higher frequency of these anchors indicates greater importance within the core content. It also includes counting the number of associations between text topic anchors and other text topic anchors; more associations indicate a wider network of connections within the text's topic anchor network and a greater role in integrating textual information. Finally, it includes counting the number of clauses in which the text topic anchor is used to explain the core meaning of the text; a higher number of clauses indicates stronger explanatory power.
[0119] The four indicators mentioned above are standardized, mapping their values to between 0 and 1 to eliminate the influence of different indicator dimensions. Then, according to preset proportions, such as semantic contribution accounting for 30%, frequency of occurrence accounting for 20%, number of associations accounting for 25%, and number of clauses accounting for 25%, the standardized indicator values are weighted and summed to obtain the semantic importance assessment result of the text topic anchor.
[0120] Step S124: Analyze the visual salience of each visual object anchor point in the scientific and technological information image, calculate the color contrast between the visual object anchor point and the surrounding visual area, calculate the difference in shape complexity between the visual object anchor point and the surrounding visual area, calculate the difference in texture density between the visual object anchor point and the surrounding visual area, standardize the visual salience, color contrast, shape complexity difference, and texture density difference respectively, and fuse the standardized visual salience, color contrast, shape complexity difference, and texture density difference according to a preset ratio to obtain the visual salience evaluation result of the visual object anchor point.
[0121] When analyzing the visual salience of anchor points, consider whether the anchor point is easily noticed in the image. Visual salience can be measured using results based on saliency detection algorithms; the higher the saliency value, the greater the visual salience.
[0122] When calculating color contrast, the color difference between the anchor point of a visual object and a certain range of surrounding visual areas is compared. For grayscale images, the difference in grayscale values is calculated; for color images, the differences in each channel in the RGB color space or the differences in hue, saturation, and brightness in the HSV color space are calculated. The greater the color contrast, the easier it is for the anchor point of the visual object to be distinguished from the background.
[0123] When calculating the difference in shape complexity, shape complexity metrics such as the tortuosity of the contour and the ratio of area to perimeter are used to compare the shape complexity of the visual object's anchor point with the surrounding visual region. The greater the difference in shape complexity, the more unique the visual object's anchor point is.
[0124] When calculating texture density differences, the density of texture features at the anchor point of the visual object and within the surrounding visual region is statistically analyzed, such as the number of texture elements per unit area. The greater the difference in texture density, the more significant the difference in texture features between the two objects.
[0125] Visual salience, color contrast, shape complexity difference, and texture density difference are standardized and mapped to a range of 0 to 1. Then, according to a preset proportion, such as visual salience accounting for 35%, color contrast accounting for 25%, shape complexity difference accounting for 20%, and texture density difference accounting for 20%, the standardized index values are weighted and summed to obtain the visual salience evaluation result of the visual object anchor point.
[0126] Step S125: Establish weight calculation rules for semantic importance and visual saliency, convert the semantic importance assessment results into semantic weight parameters, convert the visual saliency assessment results into visual weight parameters, and perform fusion calculation according to the preset ratio of semantic weight parameter proportion and visual weight parameter proportion to obtain weight fusion rules.
[0127] When establishing weight calculation rules, the semantic importance assessment results and visual saliency assessment results first need to be converted into semantic weight parameters and visual weight parameters, respectively. The conversion method can be to directly use the assessment results as weight parameters, or to perform further normalization processing.
[0128] The semantic and visual weight parameters are preset to have different percentages, such as 60% for semantic weight and 40% for visual weight. The weight fusion rule is to sum the semantic and visual weight parameters according to this ratio to obtain a comprehensive weight parameter. For example, if the semantic weight parameter is 0.8 and the visual weight parameter is 0.6, then the comprehensive weight parameter is 0.8 × 60% + 0.6 × 40% = 0.72.
[0129] Step S126: According to the weight fusion rule, the semantic weight parameters of the text topic anchor points connected by each semantic transmission path are fused and calculated with the visual weight parameters of the visual object anchor points to obtain the transmission weight parameters of the semantic transmission path.
[0130] For each semantic transmission path, i.e., the connection edge between the text topic anchor and the visual object anchor, the semantic weight parameter of the text topic anchor and the visual weight parameter of the visual object anchor are obtained. Then, according to the weight fusion rule, these two parameters are fused and calculated to obtain the transmission weight parameter of the semantic transmission path. The transmission weight parameter reflects the strength of information transmission along the path; the larger the weight parameter, the higher the priority of information transmission.
[0131] Step S127: Collect the transmission weight parameters of all semantic transmission paths, calculate the sum of all transmission weight parameters, divide each transmission weight parameter by the sum to obtain the standardized transmission weight, and make the sum of all transmission weights a fixed proportion.
[0132] After collecting the transmission weight parameters for all semantic transmission paths, their sum is calculated. Then, each transmission weight parameter is divided by the sum to obtain the standardized transmission weight. The standardized transmission weight values are between 0 and 1, and the sum of all transmission weights is 1. This process ensures that the weights of different semantic transmission paths are comparable, facilitating subsequent information transmission according to weight ratios.
[0133] Step S128: The semantic description information of the text topic anchor point is transmitted to the associated visual object anchor point along the semantic transmission path according to the corresponding transmission weight, and the semantic description information is added to the visual attribute feature description of the visual object anchor point.
[0134] Step S1281: Extract the semantic description information of the text topic anchors, and split the semantic description information into core semantic segments and supplementary explanation segments. The core semantic segments are the content used to represent the core meaning of the anchors, and the supplementary explanation segments are the content used to further explain the core semantics.
[0135] Extract the semantic description information of the text's topic anchors, i.e., their semantic description statements. Then, based on semantic logic, break down the semantic description information into core semantic segments and supplementary explanatory segments. The core semantic segment contains the most essential meaning of the anchor, such as "the learning rate parameter used to adjust the parameter update step size in gradient descent-based optimization algorithms"; the supplementary explanatory segment provides a detailed explanation of the core semantic segment, such as "the size of the learning rate parameter directly affects the convergence speed and stability of the algorithm; a larger learning rate may lead to convergence oscillations, while a smaller learning rate may lead to slow convergence."
[0136] Step S1282: Determine the range of semantic description information to be transmitted based on the proportion of transmission weight. When the proportion of transmission weight exceeds the preset weight proportion, transmit the complete content of the core semantic segment and the supplementary explanation segment; when the proportion of transmission weight does not exceed the preset weight proportion, transmit the content of the core semantic segment.
[0137] The preset weight ratio is determined based on the information delivery requirements, such as 0.5. When the transmission weight ratio between the text topic anchor and the visual object anchor exceeds 0.5, it indicates that the information transmission along this path is of high importance, and the complete content of the core semantic segment and supplementary explanatory segment needs to be transmitted so that the visual object anchor can fully understand the semantic information of the text topic anchor. When the transmission weight ratio does not exceed 0.5, only the content of the core semantic segment is transmitted to avoid excessive redundant information affecting the main attributes of the visual object anchor.
[0138] Step S1283: According to the direction of the semantic transmission path, the determined semantic description information is transmitted to the associated visual object anchor point, and the correspondence between the transmitted semantic description information and the identification information of the visual object anchor point is established.
[0139] Following the semantic transmission path, semantic description information within a defined scope is transferred from the text topic anchor to the associated visual object anchor. During this transfer, it is necessary to record the transmitted semantic description information and the corresponding visual object anchor's identification information, establishing a correspondence between the two to ensure that the information can be accurately bound to the target anchor.
[0140] Step S1284: Add a semantic association annotation region to the visual attribute feature description of the visual object anchor point. The semantic association annotation region is located at the end of the visual attribute feature description and is distinguished from the content of the visual attribute feature description.
[0141] Within the visual attribute feature description text of visual object anchor points, a dedicated semantic association annotation area is designated. This area is typically located at the end of the visual attribute feature description and is distinguished from the preceding content using specific markers or formats, such as "[Semantic Association Annotation]". This clearly separates the visual attribute feature description and semantic association annotation information, facilitating reading and subsequent processing.
[0142] Step S1285: Fill the semantic description information into the semantic association annotation area, and at the same time annotate the text topic anchor point identifier and the transmission weight ratio corresponding to the semantic description information.
[0143] The semantic description information passed from the text topic anchor will be filled into the semantic association annotation area. At the same time, in order to clarify the source and transmission strength of the information, it is necessary to annotate the text topic anchor corresponding to the semantic description information, such as "text topic anchor ID: T-001", and the transmission weight ratio of the semantic transmission path, such as "transmission weight ratio: 0.65".
[0144] Step S1286: After completing the association annotation of semantic description information, integrate the visual attribute feature description of the visual object anchor with the semantic association annotation to form a complete description of the visual object anchor carrying the semantic information of the text topic anchor.
[0145] The visual attribute feature descriptions with semantic association annotations are integrated to form a complete description of the visual object anchor. The complete description includes both the visual attribute features of the visual object anchor itself and the semantic information conveyed from the related text topic anchor, thus achieving the fusion of visual and semantic information.
[0146] Step S1287: Store information for visual object anchors carrying semantic information.
[0147] The integrated visual object anchor points are fully described and stored in a structured form, such as in the corresponding table in the database or in a local file, to ensure the security and accessibility of the information for use in the subsequent generation of cross-modal semantically enhanced representations.
[0148] Step S129: Transfer the visual attribute information of the visual object anchor point to the associated text topic anchor point along the semantic transmission path according to the corresponding transmission weight, and add the association description of the visual attribute feature in the semantic description statement of the text topic anchor point.
[0149] Similar to step S128, the visual attribute information of the visual object anchor points is first extracted, including shape features, color features, texture features, etc. The range of visual attribute information to be transmitted is determined according to the proportion of transmission weights. A higher proportion of transmission weights means more detailed visual attribute information is transmitted, while a lower proportion means more core visual attribute information is transmitted.
[0150] Then, the visual attribute information is transmitted to the associated text topic anchors along the semantic transmission path. A visual association description area is added to the semantic description statement of the text topic anchors, the transmitted visual attribute information is filled in, and the corresponding visual object anchor identifier and transmission weight ratio are marked to form a complete description of the text topic anchors carrying visual attribute information, and then stored.
[0151] Step S1210: Collect all visual object anchors carrying semantic information and all text topic anchors carrying visual information, extract the enhanced information content carried by each anchor, and remove duplicate enhanced information descriptions.
[0152] Collect all visual object anchors and text topic anchors that carry semantic information after information transmission. Extract additional enhanced information from the complete descriptions of these anchors, such as semantic association annotations for visual object anchors and visual association descriptions for text topic anchors.
[0153] The extracted augmented information content is deduplicated by comparing the augmented information carried by different anchors and removing identical or highly similar duplicate descriptions to ensure the uniqueness and conciseness of the augmented information.
[0154] Step S1211: Integrate the information of the anchors carrying enhanced information, and pair and associate the enhanced information of the text topic anchors and the visual object anchors according to the anchor identification information to form a cross-modal semantic enhanced representation containing bidirectional information transmission.
[0155] Based on the anchor point's identification information, the text topic anchor point is paired with the enhanced information of the associated visual object anchor point. For example, if the text topic anchor point T-001 is associated with the visual object anchor point V-001, then the visual association description information of T-001 is paired with the semantic association annotation information of V-001.
[0156] All paired enhancement information is integrated to form a cross-modal semantic enhancement representation. In this representation, each text topic anchor and visual object anchor contains enhancement information from the other, achieving deep fusion of text and image modal information.
[0157] Step S1210: Collect all visual object anchors carrying semantic information and all text topic anchors carrying visual information, extract the enhanced information content carried by each anchor, and remove duplicate enhanced information descriptions.
[0158] Collect all visual object anchors and text topic anchors that carry semantic information after information transmission. Extract additional enhanced information from the complete descriptions of these anchors, such as semantic association annotations for visual object anchors and visual association descriptions for text topic anchors.
[0159] The extracted augmented information content is deduplicated by comparing the augmented information carried by different anchors and removing identical or highly similar duplicate descriptions to ensure the uniqueness and conciseness of the augmented information.
[0160] Step S1211: Integrate the information of the anchors carrying enhanced information, and pair and associate the enhanced information of the text topic anchors and the visual object anchors according to the anchor identification information to form a cross-modal semantic enhanced representation containing bidirectional information transmission.
[0161] Based on the anchor point's identification information, the text topic anchor point is paired with the enhanced information of the associated visual object anchor point. For example, if the text topic anchor point T-001 is associated with the visual object anchor point V-001, then the visual association description information of T-001 is paired with the semantic association annotation information of V-001.
[0162] All paired enhancement information is integrated to form a cross-modal semantic enhancement representation. In this representation, each text topic anchor and visual object anchor contains enhancement information from the other, achieving deep fusion of text and image modal information.
[0163] Step S130: Perform hierarchical semantic parsing on the cross-modal semantic enhancement representation. At the topic level, identify the association patterns between different text topic anchors to form topic association rules. At the technology level, analyze the technology element dependency methods corresponding to technology-related anchors to form technology element dependency relationships. At the evolution level, track the performance changes of the same anchor in different scientific and technological intelligence fragments to form a concept evolution sequence. Integrate the topic association rules, technology element dependency relationships, and concept evolution sequences to obtain the scientific and technological intelligence parsing conclusion.
[0164] Layered semantic parsing of cross-modal semantic enhancement representations is a process of deeply mining scientific and technological intelligence content from different dimensions. The topic layer focuses on the correlation patterns between text topic anchors, the technology layer analyzes the technical element dependencies of technology-related anchors, the evolution layer tracks the dynamic changes of anchors, and finally, the analysis results of these three layers are integrated to form a comprehensive analytical conclusion on scientific and technological intelligence.
[0165] Step S131: Perform topic-level parsing on the cross-modal semantic enhancement representation, extract all text topic anchors and their carried semantic enhancement information, classify them according to the sub-topic categories to which the text topic anchors belong, and group text topic anchors of the same sub-topic category together.
[0166] All text topic anchors are extracted from the cross-modal semantic augmentation representation. These anchors carry semantic augmentation information passed from the visual object anchors. The text topic anchors are categorized according to the sub-topic category they belong to in their semantic descriptions, such as "optimization algorithm principles," "experimental data collection," and "performance evaluation metrics." Text topic anchors belonging to the same sub-topic category are grouped together to facilitate the analysis of relationships between anchors within the same group.
[0167] Step S132: Analyze the semantic associations between text topic anchors within each group, compare the keyword similarity clauses in the semantic enhancement information, compare the sentence logical relationship clauses in the semantic enhancement information, and identify various association patterns such as parallel association, subordinate association, and causal association between anchors.
[0168] For text topic anchor groups within the same subtopic category, analyze their semantic relationships. Compare the keyword similarity in semantic enhancement information; the more similar the keywords, the stronger the relationship. Compare the logical relationships between sentences, such as "because...therefore..." indicating a causal relationship, "at the same time...in addition..." indicating a parallel relationship, and "belonging to...includes..." indicating a subordinate relationship, etc.
[0169] These comparisons identify the association patterns between textual topic anchors. Parallel associations refer to anchors of equal status, jointly describing different aspects of a subtopic; subordinate associations refer to one anchor being a component or specific instance of another anchor; causal associations refer to one anchor being the cause or result of another anchor.
[0170] Step S133: Summarize the identified association patterns and combine text topic anchors with the same association type to form topic association rules. Each topic association rule records the associated text topic anchor identifier, association type, and association basis clause.
[0171] The identified association patterns are summarized and organized. For text topic anchor point combinations with the same association type, such as multiple parallel anchor point pairs, topic association rules are formed separately. Each topic association rule should include a unique identifier for the associated text topic anchor point, the association type (parallel, subordinate, causal, etc.), and the association basis clause. The association basis clause is the specific semantic information supporting the association pattern, such as high keyword similarity and clear logical relationship between statements.
[0172] Step S134: Perform technical layer parsing on the cross-modal semantic enhancement representation, and filter out text topic anchors and visual object anchors related to the technical content of scientific and technological intelligence. The enhancement information of the text topic anchors and visual object anchors related to the technical content of scientific and technological intelligence includes technical terms, technical structure descriptions or technical function descriptions.
[0173] When filtering for technology-related anchors, examine the enhanced information of textual topic anchors and visual object anchors. If the enhanced information contains technical terms such as "convolutional neural network," "backpropagation," or "feature extraction"; descriptions of technical structures such as "the network consists of an input layer, hidden layers, and an output layer" or "the algorithm includes a data preprocessing module and a model training module"; or descriptions of technical functions such as "used to improve the classification accuracy of the model" or "to achieve dimensionality reduction of data," then these are considered anchors related to the technical content of scientific and technological intelligence.
[0174] Step S135: Extract the technical elements corresponding to the technical anchor points, the technical concept elements corresponding to the text theme anchor points, and the technical structure elements corresponding to the visual object anchor points. Record the core content and functional terms of each technical element.
[0175] For text anchor points related to technology, extract their corresponding content as technical concept elements. Technical concept elements are abstract technical concepts, principles, methods, etc., such as "feature fusion method based on attention mechanism." Record the core content of the technical concept element, namely the definition and principles of the concept; and the functional clauses, that is, the role of the concept in the technical solution.
[0176] For technology-related visual object anchors, their corresponding content is extracted as technical structural elements. Technical structural elements are specific technical entities, components, data representations, etc., such as "structural diagram of the feature fusion module" or "experimental data visualization curve." The core content of each technical structural element is recorded, namely its composition and form; its functional role is also recorded, i.e., its function within the technical solution, such as "intuitively displaying the feature fusion effect."
[0177] Step S136: Analyze the dependency relationship between technical concept elements and technical structural elements, determine whether the technical structural elements are supporting elements required to realize the technical concept elements, determine whether the technical concept elements are functional description elements of the technical structural elements, and determine the dependency direction clauses and dependency effect clauses.
[0178] Analyzing the dependencies between technical conceptual elements and technical structural elements begins by determining whether the technical structural elements are necessary support for realizing the technical conceptual elements. For example, the technical conceptual element "feature fusion method based on attention mechanism" may require the technical structural element "attention weight calculation module" as support; without this module, the feature fusion method cannot be implemented.
[0179] Then determine whether the technical concept element is a functional description of the technical structure element. For example, the function of the technical structure element "experimental data visualization curve" is described by the technical concept element "model performance evaluation index," and the change of the curve reflects the quality of the performance index.
[0180] Based on the judgment results, determine the dependent direction clauses, such as "technical structural elements depend on technical concept elements" or "technical concept elements depend on technical structural elements"; and the dependent function clauses, such as "technical structural elements provide the basis for the realization of technical concept elements" or "technical concept elements give functional meaning to technical structural elements".
[0181] Step S137: Organize the dependency direction clauses, dependency function clauses and dependency basis clauses between technical elements to form technical element dependency relationships. Each technical element dependency relationship corresponds to a set of related technical concept elements and technical structure elements.
[0182] The dependency clauses, dependency function clauses, and supporting clauses (such as descriptions in technical documents, experimental verification results, etc.) between technical concept elements and technical structural elements are compiled together to form technical element dependency relationships. Each technical element dependency relationship clearly indicates the associated technical concept elements and technical structural elements, as well as the manner and basis of their dependency.
[0183] Step S138: Perform evolutionary layer parsing on the cross-modal semantic enhancement representation, and divide the scientific and technological intelligence into multiple scientific and technological intelligence fragments according to the time sequence of acquisition or the logical progression of content. Each scientific and technological intelligence fragment corresponds to a cross-modal semantic enhancement representation sub-part.
[0184] When segmenting scientific and technological information, if the information is a series of documents with a time sequence, such as research papers on the same technology published in different years, it should be segmented according to chronological order. If the information is a structurally complete paper, it can be segmented according to the logical progression of the content, such as "problem statement - method design - experimental verification - conclusion discussion". Each segment of scientific and technological information corresponds to a sub-part of cross-modal semantic enhancement representation, including the text topic anchors, visual object anchors, and their relationships within the segment.
[0185] Step S139: Assign a unique fragment identifier to each cross-modal semantic enhancement representation sub-part, and record the time information or logical order information corresponding to the fragment identifier.
[0186] Each cross-modal semantic enhancement representation sub-part is assigned a unique fragment identifier, such as "F-001" or "F-002". At the same time, the time information corresponding to the fragment identifier is recorded, such as "published in 2020" or "published in 2021"; or logical order information, such as "problem statement stage" or "method design stage", so as to track the evolution order of anchor points later.
[0187] Step S1310: In each cross-modal semantic enhancement representation sub-part, locate text topic anchors or visual object anchors with the same identifier based on the unique identifier information of the anchors, and track the performance of the same anchor in different scientific and technological intelligence fragments.
[0188] In each cross-modal semantic enhancement representation sub-part, the unique identifier information of the anchor, such as "T-001" or "V-001", is used to find text topic anchors or visual object anchors with the same identifier. The anchors with the same identifier represent the performance of the same concept or object in different scientific and technological intelligence fragments. By tracking them, the changes of the concept or object can be analyzed.
[0189] Step S1311: Compare the differences in semantic description statements of the same identifier text topic anchor point in different cross-modal semantic enhancement representation sub-parts, record the addition, deletion or replacement of keywords in the semantic description statements, and record the adjustment of statement structure in the semantic description statements.
[0190] For text topic anchors with the same identifier, compare their semantic description statements in different cross-modal semantically enhanced representation sub-parts. Analyze changes in keywords within the statements, such as the addition of the keyword "adaptive," the deletion of the keyword "fixed," and the replacement of "gradient descent" with "stochastic gradient descent," and record these additions, deletions, and replacements. Simultaneously, analyze adjustments to the statement structure, such as changes in word order and the addition or removal of modifiers, and record the statement structure adjustment clauses.
[0191] Step S1312: Compare the differences in visual attribute features of the same visual object anchor point in different cross-modal semantic enhancement representation sub-parts, record the changes in boundary contour coordinates in shape features, the adjustments in grayscale change data in texture features, and the changes in the distribution ratio of RGB color values in color features.
[0192] For the same visual object anchor point, compare its visual attribute features in different sub-parts. In terms of shape features, record changes in boundary contour coordinates, such as the movement of the inflection point coordinates of a curve; in terms of texture features, record adjustments to grayscale change data, such as an increase or decrease in the frequency of grayscale changes; in terms of color features, record changes in the distribution ratio of RGB color values, such as the proportion of the red channel changing from 30% to 40%, and form corresponding change clauses for each.
[0193] Step S1313: Arrange the semantic description changes of the same anchor point and the visual attribute feature changes of the same anchor point in sequence according to the time or logical order of the scientific and technological information fragments, and label the cross-modal semantic enhancement representation sub-part identifier corresponding to each change.
[0194] Arrange the clauses with changes in semantic descriptions and visual attributes that share the same anchor point according to the chronological or logical order of the scientific and technological information fragments. For example, arrange the clauses in the order of "F-001→F-002→F-003", and mark the corresponding sub-parts after each clause, such as "(F-001→F-002)", to clearly show the process and stages of the change.
[0195] Step S1314: Process the combination of semantic description changes and visual attribute changes of the same anchor point arranged according to the time sequence or logical sequence of the scientific and technological information fragments, remove minor changes that do not affect the core meaning and representation of the anchor point, and retain key changes that can reflect the changes in the core meaning or representation of the anchor point to form a concept evolution sequence.
[0196] The rearranged combinations of changes are then filtered, removing minor changes such as the addition or removal of function words, punctuation changes, or insignificant coordinate adjustments in visual attributes. These changes do not affect the core meaning and representation of the anchor. Key changes are retained, such as keyword replacements, alterations to core semantics, and significant changes to visual attributes. These changes reflect the evolution of the anchor's core meaning or representation. The retained key changes are then arranged sequentially to form a conceptual evolution sequence.
[0197] Step S1315: Collect topic association rules, technical element dependencies, and concept evolution sequences. Perform a logical consistency check on the topic association rules, technical element dependencies, and concept evolution sequences. Check whether there are any contradictions between the anchor point association clauses in the topic association rules and the element association clauses in the technical element dependencies. Check whether the technology development clauses in the technical element dependencies match the change trend clauses in the concept evolution sequences.
[0198] Collect the theme association rules, technological element dependencies, and concept evolution sequences obtained from the theme layer, technology layer, and evolution layer. Perform a logical consistency check to see if the relationships between anchor points in the theme association rules conflict with the relationships between technological elements in the technological element dependencies. For example, if A and B are parallel in the theme association rules, but A depends on B in the technological element dependencies, there is a contradiction. Check whether the technological development direction described in the technological element dependencies is consistent with the changing trend of anchor points in the concept evolution sequence. For example, if the technological element dependencies mention that technology is developing towards "intelligentization," but the changes in anchor points in the concept evolution sequence do not reflect this trend, there is a mismatch.
[0199] Step S1316: Remove duplicate related information clauses, duplicate dependency description clauses, and duplicate evolution node clauses. Integrate the topic association rules, technical element dependencies, and concept evolution sequences according to the content structure of scientific and technological intelligence to form logically coherent and content-based scientific and technological intelligence analysis conclusions.
[0200] The topic association rules, technical element dependencies, and concept evolution sequences are deduplicated to remove duplicate clauses and nodes. Then, following the content structure of scientific and technological intelligence, such as "research background - technical principles - experimental verification - development trends," these three elements are integrated. This ensures that the integrated analytical conclusions are logically coherent, comprehensive, and systematically present in-depth analytical results of the scientific and technological intelligence.
[0201] Step S140: Reverse map the topic association rules in the science and technology intelligence analysis conclusion to the cross-modal semantic anchor set. Adjust the association mapping strength between the corresponding text topic anchor and the visual object anchor based on the number of association basis clauses, the proportion of paragraphs where the association occurs, and the number of sub-topics covered by the association. Eliminate associations that exceed the reasonable association range through multiple rounds of strength calibration to obtain the updated cross-modal semantic anchor set.
[0202] The topic association rules from the analysis of scientific and technological intelligence are back-mapped to a cross-modal semantic anchor set. The aim is to optimize the association strength between anchors based on the association rules obtained from the analysis. By analyzing factors such as the number of association basis clauses, the proportion of paragraphs with associations, and the number of subtopics covered by associations, the strength parameters of the association mapping relationship are adjusted and multiple rounds of calibration are performed to ensure that the association strength is within a reasonable range, thereby obtaining an updated cross-modal semantic anchor set.
[0203] Step S141: Deconstruct the topic association rules in the science and technology intelligence analysis conclusion, extract the text topic anchor mark and visual object anchor mark involved in each topic association rule, and determine the anchor point pairs that need to be adjusted in terms of association strength.
[0204] By analyzing the topic association rules in the science and technology intelligence analysis results, each rule involves two or more text topic anchors. Through the association mapping relationships in the cross-modal semantic anchor set, the visual object anchor identifiers associated with these text topic anchors are found. This identifies the text topic anchor and visual object anchor pairs whose association strength needs to be adjusted.
[0205] Step S142: Locate the extracted text topic anchors and extracted visual object anchors in the cross-modal semantic anchor set based on the anchor identifiers, and find the original association mapping relationship and corresponding association strength parameters between the extracted text topic anchors and extracted visual object anchors.
[0206] Based on the extracted text topic anchor tags and visual object anchor tags, the corresponding anchors are found in the cross-modal semantic anchor set. Then, the original association mapping relationship records between these two anchors are searched to obtain the corresponding association strength parameter, which reflects the degree of association between the two before optimization.
[0207] Step S143: Analyze the number of related clauses for each topic association rule, count the percentage of paragraphs in which the association rule appears in the scientific and technological intelligence text, count the number of sub-topics covered by the association rule, standardize the number of related clauses, paragraph percentage, and number of sub-topics respectively, and merge the standardized number of related clauses, paragraph percentage, and number of sub-topics according to a preset ratio to obtain the basis for adjusting the association strength.
[0208] Analyze the number of supporting clauses for the topic association rules; the more supporting clauses, the more robust the support for the association rule. Statistically analyze the percentage of paragraphs in the scientific and technological intelligence texts where the association rule appears; a higher percentage indicates broader associations. Finally, analyze the number of sub-topics covered by the association rule; more sub-topics covered indicate more comprehensive associations.
[0209] These three indicators are standardized and mapped to a range of 0 to 1. Following a preset ratio, such as 50% for the number of related clauses, 30% for paragraphs, and 20% for subtopics, the standardized indicators are merged to obtain the basis for adjusting the association strength. This basis determines the direction and magnitude of the adjustment to the association strength parameter.
[0210] Step S144: Based on the correlation strength adjustment criteria, adjust the strength parameters of the correlation mapping relationship between the corresponding text topic anchor point and the visual object anchor point, collect all adjusted correlation strength parameters, set a reasonable range standard for correlation strength, and filter out the anchor point pairs corresponding to correlation strength parameters that exceed the range standard.
[0211] Based on the correlation strength adjustment criteria, adjust the original correlation strength parameters. Adjustment methods may include multiplying the original parameters by the adjustment criteria, or adding the difference between the adjustment criteria and the baseline value. After adjustment, collect the correlation strength parameters for all anchor point pairs.
[0212] Set a reasonable range for the association strength, such as 0.3 to 0.8. Compare the adjusted association strength parameter with this range and filter out anchor pairs with parameter values below 0.3 or above 0.8, as these anchor pairs may have abnormal association strength.
[0213] Step S145: For anchor pairs whose association strength parameters exceed the range standard, review the corresponding topic association rules to verify the accuracy of the statistics on the number of association basis clauses, the percentage of paragraphs with association occurrences, and the number of subtopics covered by association. If not, recalculate and adjust the association strength parameters. If the association strength still exceeds the range standard, analyze the association status of the anchor pair in the cross-modal semantic enhancement representation to determine whether the association strength is abnormal due to missing association clauses in the enhancement information. If there are missing association clauses in the enhancement information, supplement and improve the association clauses in the enhancement information and recalculate the association strength parameters.
[0214] For anchor pairs exceeding a reasonable range, first check for errors in the statistics of the number of related clauses, paragraph proportions, and number of subtopics. If the statistics are incorrect, recalculate and adjust the association strength parameters accordingly. If the statistics are correct but the association strength is still abnormal, check the enhancement information for that anchor pair in the cross-modal semantic enhancement representation to see if any related clauses are missing, such as the semantic enhancement information for text topic anchors not including content related to visual object anchors. If missing clauses exist, supplement the enhancement information and related clauses, and then recalculate the association strength parameters.
[0215] Step S146: Repeat the above abnormal calibration steps until the association strength parameters of all anchor point pairs are within a reasonable range, and obtain the calibrated association mapping relationship.
[0216] Repeat the abnormal calibration process of step S145 multiple times, repeatedly adjust and verify the correlation strength parameters until the correlation strength parameters of all anchor point pairs are within the preset reasonable range, ensuring the accuracy and rationality of the correlation mapping relationship, and obtain the calibrated correlation mapping relationship.
[0217] Step S147: Integrate the calibrated association mapping relationship with the original text topic anchors and the original visual object anchors, and replace the association mapping relationship part in the original cross-modal semantic anchor set to form an updated cross-modal semantic anchor set.
[0218] The calibrated association mappings are integrated with the existing text topic anchors and visual object anchors in the cross-modal semantic anchor set, and the new association mappings replace some of the original ones. In this way, the updated cross-modal semantic anchor set contains optimized association strength parameters, which can more accurately reflect the association between text topic anchors and visual object anchors.
[0219] Step S150: Based on the updated cross-modal semantic anchor set, the topic association rules, technical element dependencies, and concept evolution sequences are mapped to different module positions in the cross-modal semantic anchor set. The logical connection between modules is achieved through the association mapping relationship between anchors. Add association description statements between modules to generate a structured science and technology intelligence analysis report.
[0220] Based on the updated cross-modal semantic anchor set, the parsed topic association rules, technical element dependencies, and concept evolution sequences are organized into different modules. Logical connections between modules are achieved through the association mapping relationship between anchors, and association explanatory statements are added to finally generate a structured science and technology intelligence analysis report, making the analysis results clearer and easier to understand.
[0221] For example, step S151: Based on the updated cross-modal semantic anchor set as the basic framework, the basic framework is divided into multiple sub-frame regions according to the sub-topic categories of the anchors, and each sub-frame region corresponds to a science and technology intelligence sub-topic.
[0222] The updated cross-modal semantic anchor set serves as the basic framework for report generation. Based on the sub-topic categories of the anchors, such as "optimization algorithm principles," "experimental data collection," and "performance evaluation methods," the basic framework is divided into multiple sub-framework regions. Each sub-framework region corresponds to a scientific and technological intelligence sub-topic, containing textual topic anchors, visual object anchors, and their relationships within that sub-topic.
[0223] Step S152: According to the sub-topic category to which the topic association rules belong, map them to the text topic anchor point association positions in the corresponding sub-frame area of the basic framework, clarify the specific position of the text topic anchor point involved in each topic association rule in the sub-frame area, and form a topic association module.
[0224] Based on the sub-topic category to which the topic association rule belongs, place it in the corresponding sub-frame area within the basic framework. Within each sub-frame area, specify the exact location of the text topic anchor points involved in the topic association rule. For example, in the "Optimization Algorithm Principles" sub-frame area, place the topic association rule "T-001 and T-002 are causally related" at the association position between the anchor points T-001 and T-002. Organize all topic association rules in this way to form a topic association module.
[0225] Step S153: Add association annotations in the topic association module, annotating the association type, association basis clauses, and corresponding association strength parameters for each topic association rule.
[0226] In the topic association module, add association annotations to each topic association rule. The annotation content includes the association type, such as "causal association" or "parallel association"; the association basis clause, such as "based on keyword similarity in semantic enhancement information"; and the corresponding association strength parameter, such as "0.75". The above annotations make the information of the topic association rules richer and clearer.
[0227] Step S154: Classify the technical element dependencies according to the technical field, and map them to the technical anchor points in the sub-frame area related to the technical field in the basic framework. Clarify the positions of the text theme anchor points corresponding to the technical concept elements and the visual object anchor points corresponding to the technical structure elements involved in each technical element dependency in the sub-frame area, thus forming a technical dependency module.
[0228] The dependencies of technical elements are categorized according to technical fields, such as "machine learning" and "computer vision." These categorized dependencies are then mapped to relevant sub-frame areas within the basic framework, clarifying the positions of the textual anchor points corresponding to technical concept elements and the visual object anchor points corresponding to technical structure elements within these sub-frame areas. For example, the dependencies of technical elements in the "machine learning field" sub-frame area are placed at the relevant anchor points, forming a technical dependency module.
[0229] Step S155: Add dependency descriptions to the technology dependency module, describing the dependency direction clauses, dependency effect clauses, and the impact clauses of the dependency on the technical content of scientific and technological intelligence for each technical element.
[0230] In the technology dependency module, add dependency descriptions for each technology element dependency relationship. The descriptions include dependency direction clauses, such as "Technology structural elements depend on technology concept elements"; dependency effect clauses, such as "Technology structural elements provide implementation support for technology concept elements"; and dependency impact clauses on the content of scientific and technological intelligence, such as "This dependency relationship ensures the feasibility and effectiveness of the technical solution".
[0231] Step S156: According to the anchor point identification information, the concept evolution sequence is mapped to the sub-frame area where the same anchor point is located in the basic framework. According to the time order or logical order of the scientific and technological information fragments, the change clauses of the anchor point are arranged next to the anchor point to form a concept evolution module.
[0232] Based on the anchor point's identification information, the changing clauses in the concept evolution sequence are mapped to the sub-frame area containing the same anchor point in the basic framework. Next to that anchor point, the changing clauses are arranged sequentially according to the chronological or logical order of the scientific and technological information fragments, such as "2020: Keyword 'fixed learning rate'; 2021: Keyword 'adaptive learning rate' (replacement clause)". The concept evolution sequences of all anchor points are then organized to form a concept evolution module.
[0233] Step S157: Add evolution annotations in the concept evolution module, annotating the cross-modal semantic enhancement representation sub-part identifier, change content clause, and change reason clause corresponding to each change clause.
[0234] In the concept evolution module, add evolution annotations to each change clause. The annotation content includes the identifier of the cross-modal semantic enhancement representation sub-part corresponding to the change clause, such as "F-001→F-002"; the change content clause, such as "keyword replacement: fixed learning rate → adaptive learning rate"; and the change reason clause, such as "to improve the convergence speed and stability of the algorithm".
[0235] Step S158: Analyze the logical relationships between the topic association module, the technology dependency module, and the concept evolution module. Through the association mapping relationship in the updated cross-modal semantic anchor set, determine the connection points between modules. The connection points include the text topic anchor in the topic association module being connected to the visual object anchor in the technology dependency module through the association mapping relationship, and the technical elements in the technology dependency module being connected to the same anchor in the concept evolution module through the association mapping relationship.
[0236] The logical relationships between the three modules are analyzed: the topic association module describes the associations between text topic anchors, the technology dependency module describes the dependencies between technology elements, and the concept evolution module describes the changes in anchors. Connection points between modules are found through the association mapping relationships in the updated cross-modal semantic anchor set. For example, the text topic anchor T-001 in the topic association module is associated with the visual object anchor V-001 in the technology dependency module through an association mapping relationship, forming a connection point; the technology elements (T-001 and V-001) in the technology dependency module are associated with the same anchor T-001 / V-001 in the concept evolution module through an association mapping relationship, forming another connection point.
[0237] Step S159: Add association description statements between modules. The association description statements are used to describe how the topic association rules in the topic association module provide semantic support for the technical element dependencies in the technology dependency module through anchor point association clauses, and to describe how the technical element dependencies in the technology dependency module affect the concept evolution sequence in the concept evolution module through anchor point change clauses.
[0238] Add explanatory statements regarding the relationships between modules, explaining how the topic-related module supports the technology-dependent module. For example, "The causal relationship rules between T-001 and T-002 in the topic-related module, through anchor point association clauses, provide semantic support for the dependency relationship between the technical concept element corresponding to T-001 and the technical structural element corresponding to V-001 in the technology-dependent module, clarifying the causal connection between the two in function." Describe how the technology-dependent module affects the concept evolution module, such as, "Changes in the dependency relationships of technical elements in the technology-dependent module, through anchor point change clauses, affect the changing trend of the same anchor point in the concept evolution module, reflecting the changes in the connotation of the concept due to technological development."
[0239] Step S1510: Arrange the structure of the topic association module, technology dependence module, concept evolution module and related explanatory statements. Arrange the module positions according to the overall logical order of the scientific and technological information, with the topic association module first, the technology dependence module in the middle, the concept evolution module last, and the related explanatory statements interspersed between the modules.
[0240] Following the overall logical order of the scientific and technological intelligence content, the three modules and related explanatory statements are structurally arranged. Typically, the topic-related module is introduced first, allowing readers to understand the connections between core concepts; then the technology dependency module explains the dependencies in technology implementation; and finally, the concept evolution module showcases the development and changes of the concepts. Related explanatory statements are interspersed between the modules, serving as connections and transitions, making the overall report structure clearer and more coherent.
[0241] Step S1511: Add a cover, table of contents, and summary to the science and technology intelligence analysis report. The cover includes the report name, analysis object, and analysis time. The table of contents includes page number indexes for the topic association module, technology dependence module, concept evolution module, and related explanatory statements. The summary briefly summarizes the core content clauses of the science and technology intelligence analysis conclusions.
[0242] Add a cover page to the report, indicating the report title, such as "In-depth Analysis Report on Deep Learning Model Optimization Technology"; the analysis object, such as "A series of academic papers and experimental charts on deep learning model optimization technology"; and the analysis time, such as "October 2023".
[0243] Compile a table of contents, listing page numbers for the topic association module, technology dependency module, concept evolution module, and related explanatory statements within the report, facilitating reader access. Write an abstract, concisely summarizing the core content of the scientific and technological intelligence analysis conclusions, such as the main topic association rules, key technological element dependencies, and important concept evolution trends.
[0244] Step S1512: Standardize the format of the report content to form a structured scientific and technological intelligence analysis report.
[0245] The font, font size, line spacing, paragraph formatting, and heading levels of the report are standardized to ensure a consistent and aesthetically pleasing format. The content of each module, related explanatory statements, cover page, table of contents, and abstract are integrated to form a complete and structured scientific and technological intelligence analysis report.
[0246] Based on the same inventive concept, please refer to Figure 2 The diagram shows a schematic block diagram of a deep analysis system 100 based on cross-modal semantic enhancement for performing the above-described deep analysis method for scientific and technological intelligence based on cross-modal semantic enhancement, provided in an embodiment of this application. The deep analysis system 100 based on cross-modal semantic enhancement may include a communication unit 110, a machine-readable storage medium 120, and a processor 130.
[0247] Alternatively, the machine-readable storage medium 120 can also be integrated into the processor 130 and can communicate and interact with external systems through the communication unit 110. The machine-readable storage medium 120 stores machine-executable instructions for executing the scheme of this application, and the processor 130 executes the machine-executable instructions stored in the machine-readable storage medium 120 to implement the deep analysis method for scientific and technological intelligence based on cross-modal semantic enhancement provided in the aforementioned method embodiments.
[0248] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.
Claims
1. A technology intelligence deep analysis method based on cross-modal semantic enhancement, characterized in that, The method comprises: constructing a cross-modal semantic anchor point set, the cross-modal semantic anchor point set containing text theme anchors extracted from scientific intelligence texts, visual object anchors extracted from scientific intelligence images, and an association mapping relationship between the text theme anchors and the visual object anchors, the text theme anchors carrying semantic description information representing core themes of the scientific intelligence, and the visual object anchors carrying visual attribute information representing key objects of the scientific intelligence; based on the cross-modal semantic anchor point set, constructing a semantic conduction path between anchors, realizing bidirectional information transmission between the text theme anchors and the visual object anchors through the semantic conduction path, injecting semantic description information of the text theme anchors into the visual object anchors, and simultaneously injecting visual attribute information of the visual object anchors into the text theme anchors, to generate a cross-modal semantic enhanced representation containing bidirectional enhanced information; performing hierarchical semantic analysis on the cross-modal semantic enhanced representation, identifying an association mode between different text theme anchors at a theme layer to form a theme association rule, analyzing a technical element dependency mode of technical related anchors at a technology layer to form a technical element dependency relationship, and tracking a change in performance of a same anchor in different scientific intelligence segments at an evolution layer to form a concept evolution sequence, and integrating the theme association rule, the technical element dependency relationship, and the concept evolution sequence to obtain a scientific intelligence analysis conclusion; mapping the theme association rule in the scientific intelligence analysis conclusion to the cross-modal semantic anchor point set in a reverse direction, adjusting an association mapping relationship strength between corresponding text theme anchors and visual object anchors according to a number of association basis clauses of the theme association rule, a paragraph proportion of an association occurrence, and a number of sub-themes covered by the association, eliminating associations beyond a reasonable association range through multiple rounds of strength calibration, and obtaining an updated cross-modal semantic anchor point set; based on the updated cross-modal semantic anchor point set, corresponding the theme association rule, the technical element dependency relationship, and the concept evolution sequence to different module positions of the cross-modal semantic anchor point set, realizing logical connection of the modules through the association mapping relationship between the anchors, adding association description sentences between the modules, and generating a structured scientific intelligence analysis report.
2. The sci-tech intelligence deep analysis method based on cross-modal semantic enhancement according to claim 1, characterized in that, The method comprises: reading complete content of the scientific intelligence texts, splitting the scientific intelligence texts into multiple text paragraph units according to semantic logical relationships between paragraphs, and each text paragraph unit corresponding to a sub-theme content of the scientific intelligence; performing sentence splitting on each text paragraph unit, splitting the text paragraph unit into multiple independent text sentences, and removing duplicate sentences and semantically incomplete sentences containing redundant information; performing vocabulary segmentation processing on each retained text sentence to obtain multiple vocabulary units, filtering virtual vocabulary units without actual semantics and low-frequency irrelevant vocabulary units appearing only once in the scientific intelligence texts, and retaining core vocabulary units having semantic value of the scientific intelligence; combining the core vocabulary units according to semantic association to form multiple vocabulary combinations capable of representing core meanings of the sub-themes, and simultaneously retaining sentence fragments that cannot be further split and are semantically complete, the vocabulary combinations and the semantically complete sentence fragments constituting a candidate text theme anchor point set; For each candidate anchor point in the candidate text theme anchor point set, count the number of occurrences of each candidate anchor point in all paragraphs of the sci-tech intelligence text, count the number of occurrences of each candidate anchor point in key paragraphs that set forth core technology, core viewpoints, and core conclusions, count the number of keywords that coincide between each candidate anchor point and the core theme description statement of the sci-tech intelligence text, and prioritize the candidate anchor points in order from most to least occurrences, from most to least occurrences in key paragraphs, and from most to least number of keyword coincidences; Select the candidate anchor points with high priority as the final text theme anchor points, and generate corresponding semantic description statements for each text theme anchor point, which fully cover the core semantics of the text theme anchor points and the sub-theme category to which they belong; Load the complete pixel data of the sci-tech intelligence image, and use a Gaussian filter algorithm to remove visual noise from the pixel data and retain visual content that reflects key objects of the sci-tech intelligence; Use a region growing-based visual region segmentation algorithm to divide the denoised sci-tech intelligence image into regions, and segment the sci-tech intelligence image into multiple non-overlapping visual regions based on the similarity of pixel gray values, with each visual region corresponding to a potential visual object; Perform object recognition on each visual region by comparing the pixel distribution characteristics of the visual region with the pre-set sci-tech intelligence key object feature library, retaining visual regions with matching pixel distribution characteristics and any feature in the key object feature library, and filtering out visual regions containing key information objects; For visual regions containing key information objects, extract the coordinate data of the boundary contour of the visual region to obtain shape features, extract the frequency and amplitude data of the pixel gray value changes within the visual region to obtain texture features, and extract the RGB color value distribution proportion data of each pixel point within the visual region to obtain color features; Determine the visual regions with extracted shape features, texture features, and color features as visual object anchor points, generate corresponding visual attribute feature descriptions for each visual object anchor point, and record the coordinate position information of each visual object anchor point in the sci-tech intelligence image; Establish a mapping relationship between the semantic description statements of the text theme anchor points and the visual attribute feature descriptions of the visual object anchor points, calculate the association similarity by counting the number of keywords that coincide between the semantic description statements and the visual attribute feature descriptions, counting the number of corresponding clauses in the pre-set semantic-visual association dictionary, and counting the number of matching clauses in overall semantic comparison, and pair text theme anchor points and visual object anchor points with an association similarity that exceeds a pre-set association standard; Perform cross-modal semantic consistency verification on the paired anchor point combinations, and eliminate anchor point combinations with different semantic meanings and visual representations pointing to different sci-tech intelligence content. The verified text theme anchor point, the verified visual object anchor point and the mapping relationship between the verified text theme anchor point and the visual object anchor point are integrated to form a cross-modal semantic anchor point set, and each anchor point in the cross-modal semantic anchor point set carries unique identification information to distinguish different anchor points.
3. The sci-tech intelligence deep analysis method based on cross-modal semantic enhancement according to claim 2, characterized in that, The occurrence number of each candidate anchor point in all paragraphs of the sci-tech information text, the occurrence number of each candidate anchor point in key paragraphs describing core technology, core viewpoint and core conclusion, and the number of keywords of the sci-tech information text core theme description statement coinciding with each candidate anchor point are counted, and the candidate anchor points are prioritized in the order of the occurrence number from more to less, the key paragraph occurrence number from more to less and the keyword coincidence number from more to less, including: All text paragraph units of the sci-tech information text are traversed, the occurrence number of each candidate anchor point in each text paragraph unit is counted, and the occurrence number in all text paragraph units is accumulated to obtain the total occurrence number of each candidate anchor point; Key paragraphs in the sci-tech information text are identified, including paragraphs describing core technology, paragraphs describing core viewpoint and paragraphs describing core conclusion, the occurrence number of each candidate anchor point in each key paragraph is counted, and the occurrence number in all key paragraphs is accumulated to obtain the key paragraph total occurrence number of each candidate anchor point; The core theme description statement of the sci-tech information text is extracted, the core theme description statement is located in the introduction part, the abstract part or the conclusion part of the text, the core theme description statement is used to summarize the overall core content of the sci-tech information, the keywords in the core theme description statement are extracted to form a core keyword list; The keywords in the semantic description of each candidate anchor point are extracted to form a candidate keyword list, and the number of coinciding keywords of the candidate keyword list and the core keyword list is counted; The total occurrence number, the key paragraph total occurrence number and the number of coinciding keywords are used as sorting indexes, and the candidate anchor points are preliminarily sorted in the order of the total occurrence number from more to less; If the total occurrence numbers of two candidate anchor points are the same, the key paragraph total occurrence numbers of the two candidate anchor points are compared, and the candidate anchor point with the larger key paragraph total occurrence number is arranged in front; If the total occurrence numbers and the key paragraph total occurrence numbers of two candidate anchor points are the same, the number of coinciding keywords of the two candidate anchor points is compared, and the candidate anchor point with the more number of coinciding keywords is arranged in front; If the total occurrence numbers, the key paragraph total occurrence numbers and the number of coinciding keywords of two candidate anchor points are the same, the semantic description length of the candidate anchor points is compared, and the candidate anchor point with the shortest semantic description length is arranged in front; A priority sorting table of the candidate anchor points is generated, and the priority sorting table includes the identification of the candidate anchor points, the total occurrence number of the candidate anchor points, the key paragraph total occurrence number of the candidate anchor points, the number of coinciding keywords of the candidate anchor points and the sorting name of the candidate anchor points.
4. The sci-tech intelligence deep analysis method based on cross-modal semantic enhancement according to claim 2, characterized in that, The mapping relationship between the semantic description sentence of the text theme anchor point and the visual attribute feature description of the visual object anchor point is established by counting the number of overlapping keywords of the semantic description sentence and the visual attribute feature description, counting the number of corresponding clauses of the two in the preset semantic-visual association dictionary, and counting the number of matching clauses of the overall semantic comparison of the two, calculating the association similarity, and pairing the text theme anchor point and the visual object anchor point whose association similarity exceeds the preset association standard, including: Extracting core keywords in the semantic description sentence of each text theme anchor point, the core keywords being nouns, verbs or adjectives used to represent the core meaning of the semantic description sentence, removing the function words and modifying words in the semantic description sentence to form a text keyword list; Extracting key feature words in the visual attribute feature description of each visual object anchor point, the key feature words being words used to represent visual attributes, the key feature words including shape feature corresponding words, color feature corresponding words, and texture feature corresponding words to form a visual keyword list; Establishing a semantic-visual association dictionary, the semantic-visual association dictionary recording the corresponding relationship clauses of the preset scientific and technical information semantic keywords and visual feature words; Counting the number of overlapping keywords of the text keyword list and the visual keyword list, the number of overlapping keywords being the total number of words existing in both the text keyword list and the visual keyword list; Based on the semantic-visual association dictionary, counting the number of clauses in which the keywords in the text keyword list and the keywords in the visual keyword list have a corresponding relationship in the semantic-visual association dictionary, the number of corresponding relationship clauses being the total number of clauses in which the text keywords and the visual keywords form a corresponding relationship; Using a semantic similarity calculation algorithm in natural language processing, performing overall semantic comparison on the semantic description sentence of the text theme anchor point and the visual attribute feature description of the visual object anchor point, counting the number of semantic matching clauses between the two, the number of semantic matching clauses being the total number of clauses in which the semantic description sentence and the visual attribute feature description have a corresponding semantic relationship; Fusing and calculating the number of overlapping keywords, the number of corresponding relationship clauses, and the number of semantic matching clauses according to a preset proportion to obtain the association similarity of the semantic description sentence and the visual attribute feature description; Setting a preset association standard for the association similarity, the preset association standard being determined according to the field characteristics and cross-modal association requirements of scientific and technical information; Pairing the text theme anchor point and the visual object anchor point whose association similarity exceeds the preset association standard, recording the identification information and the association similarity of each pair of anchor points, and forming a preliminary anchor point pairing set; De-duplicating the preliminary anchor point pairing set, if a text theme anchor point is paired with multiple visual object anchor points, or a visual object anchor point is paired with multiple text theme anchor points, retaining the pairing combination with the highest association similarity and eliminating the remaining pairing combinations, so that each anchor point corresponds to only one associated anchor point in the pairing set.
5. The sci-tech intelligence deep analysis method based on cross-modal semantic enhancement according to claim 1, characterized in that, The cross-modal semantic anchor point set is used to construct an inter-anchor semantic conduction path, and bidirectional information transmission between a text theme anchor point and a visual object anchor point is realized through the semantic conduction path, so as to inject semantic description information of the text theme anchor point into the visual object anchor point and inject visual attribute information of the visual object anchor point into the text theme anchor point, thereby generating a cross-modal semantic enhanced representation containing bidirectional enhanced information, including: An association mapping relationship between all text theme anchor points and visual object anchor points is extracted from the cross-modal semantic anchor point set, the number of visual object anchor points corresponding to each text theme anchor point is counted, the number of text theme anchor points corresponding to each visual object anchor point is counted, and a connection relationship between the anchor points is obtained; A semantic conduction path network between the anchor points is constructed by taking the text theme anchor points and the visual object anchor points as nodes and taking the association mapping relationship between the anchor points as edges, wherein each edge in the semantic conduction path network corresponds to a semantic conduction path from a text theme anchor point to a visual object anchor point or from a visual object anchor point to a text theme anchor point; The semantic contribution degree of each text theme anchor point in the sci-tech intelligence text is analyzed, the number of occurrences of the text theme anchor point in the core paragraph is counted, the number of associations between the text theme anchor point and other text theme anchor points is counted, and the number of clauses explaining the core meaning of the text by the text theme anchor point is counted. The semantic contribution degree, the number of occurrences, the number of associations, and the number of clauses are respectively standardized, and the standardized semantic contribution degree, the number of occurrences, the number of associations, and the number of clauses are fused according to a preset proportion to obtain a semantic importance evaluation result of the text theme anchor point; The visual prominence of each visual object anchor point in the sci-tech intelligence image is analyzed, the color contrast between the visual object anchor point and the surrounding visual area is calculated, the shape complexity difference between the visual object anchor point and the surrounding visual area is calculated, and the texture density difference between the visual object anchor point and the surrounding visual area is calculated. The visual prominence, the color contrast, the shape complexity difference, and the texture density difference are respectively standardized, and the standardized visual prominence, the color contrast, the shape complexity difference, and the texture density difference are fused according to a preset proportion to obtain a visual saliency evaluation result of the visual object anchor point; A weight calculation rule of the semantic importance and the visual saliency is established, the semantic importance evaluation result is converted into a semantic weight parameter, the visual saliency evaluation result is converted into a visual weight parameter, and the semantic weight parameter proportion and the visual weight parameter proportion are fused according to a preset proportion to obtain a weight fusion rule; According to the weight fusion rule, the semantic weight parameter of the text theme anchor point and the visual weight parameter of the visual object anchor point connected by each semantic conduction path are fused and calculated to obtain a conduction weight parameter of the semantic conduction path; The conduction weight parameters of all semantic conduction paths are collected, the sum of all conduction weight parameters is calculated, each conduction weight parameter is divided by the sum to obtain a standardized conduction weight, and the sum of all conduction weights is a fixed proportion. The semantic description information of the text theme anchor point is transmitted to the associated visual object anchor point along the semantic transmission path according to the corresponding transmission weight, and the semantic description information is added to the visual attribute feature description of the visual object anchor point as an associated annotation; The visual attribute feature description of the visual object anchor point is transmitted to the associated text theme anchor point along the semantic transmission path according to the corresponding transmission weight, and the visual attribute feature description is added to the semantic description sentence of the text theme anchor point as an associated explanation; All visual object anchor points carrying semantic information and all text theme anchor points carrying visual information are collected, and the enhanced information content carried by each anchor point is extracted, and the repeated enhanced information description is removed; The information of the anchor points carrying the enhanced information is integrated, the enhanced information of the text theme anchor point and the visual object anchor point is matched and associated according to the identification information of the anchor points, and a cross-modal semantic enhancement representation containing bidirectional transmission information is formed.
6. The sci-tech intelligence deep analysis method based on cross-modal semantic enhancement according to claim 5, characterized in that, The semantic description information of the text theme anchor point is transmitted to the associated visual object anchor point along the semantic transmission path according to the corresponding transmission weight, and the semantic description information is added to the visual attribute feature description of the visual object anchor point as an associated annotation, including: The semantic description information of the text theme anchor point is extracted, and the semantic description information is divided into a core semantic segment and a supplementary explanation segment, the core semantic segment is content for representing the core meaning of the anchor point, and the supplementary explanation segment is content for further explaining the core semantic; According to the proportion of the transmission weight, the range of the transmitted semantic description information is determined, when the proportion of the transmission weight exceeds a preset weight proportion, the complete content of the core semantic segment and the supplementary explanation segment is transmitted, and when the proportion of the transmission weight does not exceed the preset weight proportion, the content of the core semantic segment is transmitted; According to the direction of the semantic transmission path, the determined semantic description information is transmitted to the associated visual object anchor point, and the corresponding association between the transmitted semantic description information and the identification information of the visual object anchor point is established; A semantic association annotation area is added to the visual attribute feature description of the visual object anchor point, the semantic association annotation area is located at the end of the visual attribute feature description, and is separated from the visual attribute feature description content; The transmitted semantic description information is filled in the semantic association annotation area, and the identification of the text theme anchor point corresponding to the semantic description information and the proportion of the transmission weight are labeled; After the association annotation of the semantic description information is completed, the visual attribute feature description of the visual object anchor point and the semantic association annotation are integrated to form a complete description of the visual object anchor point carrying the semantic information of the text theme anchor point; The visual object anchor points carrying the semantic information are stored.
7. The sci-tech intelligence deep analysis method based on cross-modal semantic enhancement according to claim 1, characterized in that, The cross-modal semantic enhancement representation is subjected to hierarchical semantic analysis, the association mode between different text theme anchor points is identified in the theme layer to form a theme association rule, the dependence mode of the technical elements corresponding to the technical related anchor points is analyzed in the technical layer to form a technical element dependence relationship, the performance change of the same anchor point in different science and technology intelligence segments is tracked in the evolution layer to form a concept evolution sequence, and the theme association rule, the technical element dependence relationship and the concept evolution sequence are integrated to obtain a science and technology intelligence analysis conclusion, including: The cross-modal semantic enhancement representation is subjected to topic layer analysis, all text theme anchors and the semantic enhancement information carried thereby are extracted, and the text theme anchors are classified according to the sub-theme categories to which the text theme anchors belong, and the text theme anchors of the same sub-theme category are grouped together; The semantic association between the text theme anchors in each group is analyzed, the keyword similarity clauses in the semantic enhancement information are compared, and the sentence logical relationship clauses in the semantic enhancement information are compared to identify the parallel association, subordinate association and cause-effect association between the anchors; The identified association modes are summarized, and the text theme anchors with the same association type are combined to form theme association rules, and each theme association rule records the anchor identification, association type and association basis clauses of the associated text theme anchors; The cross-modal semantic enhancement representation is subjected to technology layer analysis, and text theme anchors related to the technical content of the scientific and technical information and visual object anchors related to the technical content of the scientific and technical information are screened out, and the enhancement information of the text theme anchors related to the technical content of the scientific and technical information and the visual object anchors related to the technical content of the scientific and technical information contains technical terms, technical structure descriptions or technical function descriptions; The technical elements corresponding to the technology-related anchors are extracted, the technical concept elements corresponding to the text theme anchors, and the technical structure elements corresponding to the visual object anchors, and the core content and function clauses of each technical element are recorded; The dependency relationship between the technical concept elements and the technical structure elements is analyzed, it is judged whether the technical structure elements are support elements required for realizing the technical concept elements, whether the technical concept elements are function description elements of the technical structure elements, and the dependency direction clauses and dependency action clauses are determined; The dependency direction clauses, dependency action clauses and dependency basis clauses between the technical elements are sorted to form technical element dependency relationships, and each technical element dependency relationship corresponds to a group of associated technical concept elements and technical structure elements; The cross-modal semantic enhancement representation is subjected to evolution layer analysis, and the scientific and technical information is divided into multiple scientific and technical information segments in the order of time or logical progression of content, and each scientific and technical information segment corresponds to a sub-part of the cross-modal semantic enhancement representation; A unique segment identifier is assigned to each sub-part of the cross-modal semantic enhancement representation, and the time information or logical order information corresponding to the segment identifier is recorded; In each sub-part of the cross-modal semantic enhancement representation, the same-identified text theme anchors or the same-identified visual object anchors are located according to the unique anchor identification information, and the same anchor in different scientific and technical information segments is tracked; The semantic description sentence differences of the same-identified text theme anchors in different sub-parts of the cross-modal semantic enhancement representation are compared, the addition, deletion or replacement clauses of the keywords in the semantic description sentences are recorded, and the adjustment clauses of the sentence structure in the semantic description sentences are recorded; The visual attribute feature differences of the same-identified visual object anchors in different sub-parts of the cross-modal semantic enhancement representation are compared, the change clauses of the boundary contour coordinates in the shape feature are recorded, the adjustment clauses of the gray scale change data in the texture feature are recorded, and the change clauses of the RGB color value distribution proportion in the color feature are recorded. The semantic description sentence changes of the same identification anchor point and the visual attribute feature changes of the same identification anchor point are arranged in sequence according to the time sequence or the logical sequence of the scientific information segments, and each change is marked with a cross-modal semantic enhancement representation subpart identification; The combination of the semantic description sentence changes and the visual attribute feature changes of the same identification anchor point arranged according to the time sequence or the logical sequence of the scientific information segments is processed, minor changes that do not affect the core meaning and representation of the anchor point are removed, key change clauses that can reflect the core meaning or representation changes of the anchor point are retained, and a concept evolution sequence is formed; The theme association rule, the technical element dependency relationship and the concept evolution sequence are collected, and logical consistency is checked. It is checked whether the anchor point association clauses in the theme association rule and the element association clauses in the technical element dependency relationship are contradictory, and whether the technical development clauses in the technical element dependency relationship and the change trend clauses in the concept evolution sequence are matched; Repeated association information clauses, repeated dependency description clauses and repeated evolution node clauses are removed, and the theme association rule, the technical element dependency relationship and the concept evolution sequence are integrated according to the content structure of the scientific information to form a scientific information analysis conclusion.
8. The sci-tech intelligence deep analysis method based on cross-modal semantic enhancement according to claim 7, characterized in that, The semantic description sentence changes of the same identification anchor point and the visual attribute feature changes of the same identification anchor point are arranged in sequence according to the time sequence or the logical sequence of the scientific information segments, and each change is marked with a cross-modal semantic enhancement representation subpart identification, forming a concept evolution sequence, including: The semantic description sentences and the visual attribute feature descriptions of the same identification anchor point in all cross-modal semantic enhancement representation subparts are collected and arranged into a change record list according to the time sequence or the logical sequence of the corresponding scientific information segments; In the change record list, the semantic description sentences and the visual attribute feature descriptions corresponding to each cross-modal semantic enhancement representation subpart are assigned corresponding segment identifications; The semantic description sentences corresponding to adjacent two segment identifications are compared, key word change clauses, sentence structure change clauses and semantic range change clauses in the semantic description sentences are identified, and the identified change clauses are recorded in the change record list; The visual attribute feature descriptions corresponding to adjacent two segment identifications are compared, boundary contour coordinate change clauses in shape features, gray value change data change clauses in texture features and RGB color value distribution proportion change clauses in color features are identified, and the identified change clauses are recorded in the change record list; The change clauses in the change record list are classified, the change clauses of the semantic description sentences are classified into a semantic change class, and the change clauses of the visual attribute feature descriptions are classified into a visual change class; The semantic change class clauses and the visual change class clauses of the same identification anchor point are alternately arranged according to the time sequence or the logical sequence of the scientific information segments, and each segment identification corresponds to a group of semantic change clauses and visual change clauses. In the combined sequence of the semantic change clauses and the visual change clauses of the same identified anchor point after arrangement, the cross-modal semantic enhancement representation sub-part identification corresponding to each change clause is marked, and the identification is located at the end of the change clause and distinguished from the change clause by a special symbol; The combined sequence of the semantic change clauses and the visual change clauses of the same identified anchor point after arrangement is subjected to integrity checking, and it is checked whether the change clause corresponding to each cross-modal semantic enhancement representation sub-part is included in the combined sequence; Based on the preset semantic importance dictionary and visual saliency standard, the change clause judged as unimportant in the combined sequence is removed; The change clause judged as important is retained as a key change clause, the key change clause is arranged in order, and a final concept evolution sequence is formed, the concept evolution sequence includes the change content of the key change clause, the segment identification corresponding to the key change clause and the change type of the key change clause.
9. The sci-tech intelligence deep analysis method based on cross-modal semantic enhancement according to claim 1, characterized in that, The theme association rule in the scientific intelligence analysis conclusion is reversely mapped to the cross-modal semantic anchor point set, the association mapping relationship strength between the corresponding text theme anchor point and the visual object anchor point is adjusted according to the number of association basis clauses of the theme association rule, the paragraph proportion of the association occurrence and the number of sub-themes covered by the association, the association beyond the reasonable association range is eliminated through multiple rounds of strength calibration, and an updated cross-modal semantic anchor point set is obtained, including: The theme association rule in the scientific intelligence analysis conclusion is disassembled, the text theme anchor point identification and the visual object anchor point identification involved in each theme association rule are extracted, and the anchor point pairs that need to adjust the association strength are determined; The extracted text theme anchor point and the extracted visual object anchor point are located in the cross-modal semantic anchor point set according to the anchor point identification, and the original association mapping relationship and the corresponding association strength parameter between the extracted text theme anchor point and the extracted visual object anchor point are found; The number of association basis clauses of each theme association rule is analyzed, the paragraph proportion of the association rule in the scientific intelligence text is counted, and the number of sub-themes covered by the association rule is counted, the number of association basis clauses, the paragraph proportion and the number of sub-themes are standardized respectively, and the standardized number of association basis clauses, the paragraph proportion and the number of sub-themes are fused according to a preset proportion to obtain an association strength adjustment basis; According to the association strength adjustment basis, the strength parameter of the association mapping relationship between the corresponding text theme anchor point and the visual object anchor point is adjusted, all adjusted association strength parameters are collected, a reasonable range standard of the association strength is set, and the anchor point pairs corresponding to the association strength parameters beyond the range standard are screened out; For the anchor pair corresponding to the association strength parameter exceeding the range standard, the corresponding subject association rule is reviewed, and the accuracy of the statistics of the number of association basis clauses, the statistics of the proportion of the associated paragraphs, and the statistics of the number of sub-topics covered by the association is verified. If not, the association strength parameter is re-calculated and adjusted. If yes, but the association strength still exceeds the range standard, the enhanced information association of the anchor pair in the cross-modal semantic enhanced representation is analyzed to determine whether the association strength is abnormal due to the lack of association clauses in the enhanced information. If there is a lack of enhanced information association clauses, the enhanced information association clauses are supplemented and improved, and the association strength parameter is re-calculated. The operation of re-checking, verifying statistics, adjusting parameters, analyzing enhanced information association, supplementing and improving enhanced information association clauses, and re-calculating the association strength parameter of the anchor pair corresponding to the association strength parameter exceeding the range standard is repeatedly performed until the association strength parameters of all anchor pairs are within the reasonable range standard, and the calibrated association mapping relationship is obtained. The calibrated association mapping relationship is integrated with the original text subject anchor point and the original visual object anchor point, replacing the association mapping relationship part in the original cross-modal semantic anchor point set to form an updated cross-modal semantic anchor point set.
10. A technology intelligence deep analysis system based on cross-modal semantic enhancement, characterized in that, It includes: a processor; a machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to execute the machine-executable instructions to perform the cross-modal semantic enhancement-based scientific and technical information deep analysis method of any one of claims 1 to 9.
Citation Information
Patent Citations
A web page parsing method and system based on artificial intelligence
CN119760650A
Artificial intelligence-based literature structured extraction method and system
CN120336416A