Science and technology intelligence collection method and system based on multi-modal data acquisition

By calculating the intelligence timeliness index TI and modal data indices dex, ma, tex, and vid and comparing them with thresholds, multimodal data is screened and classified, solving the problem of inaccurate intelligence timeliness and quality assessment in existing technologies, and achieving efficient and accurate intelligence collection and screening.

CN119917709BActive Publication Date: 2026-03-31BEIJING SCI & TECH PATENT OFFICE
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing methods for collecting scientific and technological intelligence are insufficient to accurately control the timeliness and quality of intelligence, and lack a systematic and comprehensive evaluation system, resulting in insufficient effectiveness of the collected intelligence in practical applications.

Method used

By calculating the intelligence timeliness index TI, and comparing the corresponding indices dex, ma, tex, and vid of text, image, audio, and video data with thresholds, multimodal data is filtered and classified, a set of basic intelligence thresholds is set, and basic intelligence that meets the standards is selected.

Benefits of technology

It enables precise screening and classification of multimodal data, improves the timeliness and accuracy of intelligence gathering, ensures data quality and reliability, and meets the needs of scientific research innovation and corporate decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119917709B_ABST
    Figure CN119917709B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for collecting scientific and technological intelligence based on multimodal data acquisition, relating to the field of intelligence data processing technology. The main solution involves: collecting text, images, audio, video, and basic intelligence data; calculating an intelligence timeliness index based on the basic intelligence data; and filtering intelligence data within the timeframe by comparing the index with thresholds. Text indices, image indices, audio indices, and video indices are calculated based on the text, image, audio, and video data respectively, and these indices are compared with their respective thresholds to filter and collect standard data. The system includes corresponding modules for data collection, index calculation, and filtering. This invention can comprehensively evaluate scientific and technological intelligence by integrating multimodal data, improving the accuracy and effectiveness of intelligence collection, and has significant application value in the field of scientific and technological intelligence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligence data processing technology, specifically to a method and system for collecting scientific and technological intelligence based on multimodal data acquisition. Background Technology

[0002] In the field of science and technology intelligence, with the rapid development of information technology, intelligence data is showing a significant trend of explosive growth and multimodalization. Multimodal data encompasses rich text, images, audio, and video information. How to efficiently and accurately collect valuable science and technology intelligence from massive and complex multimodal data has become a critical issue that urgently needs to be addressed.

[0003] Existing methods for collecting scientific and technological intelligence often only involve searching for keywords in the intelligence data and collecting intelligence data with the same keywords in a unified manner, lacking quality control over the intelligence data.

[0004] Therefore, existing technologies have obvious shortcomings. For example, they are difficult to accurately control the timeliness of intelligence, and cannot promptly filter out the latest and most valuable intelligence based on the dynamic changes in data. At the same time, they lack a comprehensive and scientific evaluation system, and cannot accurately judge the quality and reliability of intelligence from the overall perspective of multimodal data. This greatly reduces the effectiveness of the collected scientific and technological intelligence in practical applications, making it difficult to meet the growing demand for high-quality scientific and technological intelligence in scientific research and innovation, corporate decision-making, and other fields. Summary of the Invention

[0005] (a) Technical problems to be solved

[0006] To address the shortcomings of existing technologies, this invention provides a method and system for collecting scientific and technological intelligence based on multimodal data acquisition. By collecting basic intelligence data and calculating an intelligence timeliness index based on it, and by filtering and classifying multimodal data according to the comparison results with the intelligence timeliness index threshold, this invention solves the problem of existing technologies having simple and inaccurate assessments of intelligence timeliness. By calculating the indices corresponding to text, image, audio, and video data separately and comparing them with their respective thresholds to filter standard data, this invention solves the problem of existing technologies lacking systematic and precise quantitative indicators when filtering data, resulting in low accuracy and efficiency in intelligence collection.

[0007] (II) Technical Solution

[0008] To achieve the above objectives, the present invention provides the following technical solution: a method and system for collecting scientific and technological intelligence based on multimodal data acquisition, comprising:

[0009] Collect basic intelligence data, which includes several basic intelligence items;

[0010] Based on basic intelligence data, calculate the intelligence timeliness index TI for several basic intelligences; set an intelligence timeliness index threshold F; based on the comparison between the intelligence timeliness index TI and the intelligence timeliness index threshold F, screen and collect basic intelligence within the time limit, and classify the screened basic intelligence into text data, image data, audio data and video data.

[0011] The text index (dex) is calculated based on the text data, and the image index (ma) and average shooting time are calculated based on the image data. The audio index tex is calculated based on the audio data, and the video index vid is calculated based on the video data.

[0012] Set a basic intelligence threshold set, compare the text index (dex), image index (ma), audio index (tex), and video index (vid) with the corresponding thresholds in the basic intelligence threshold set, and filter basic intelligence that meets the criteria.

[0013] In the preferred scheme of the above-mentioned method and system for collecting scientific and technological intelligence based on multimodal data acquisition: the specific steps for calculating the intelligence timeliness index (TI) are as follows:

[0014] Basic intelligence data includes intelligence release duration t, decay rate ki, expiration threshold vt, and maximum effective period Tmax;

[0015] The intelligence timeliness index is calculated based on the intelligence release duration t, decay rate ki, expiration threshold vt, and longest effective period Tmax. The specific formula used is as follows:

[0016] .

[0017] In the preferred scheme of the above-mentioned method and system for collecting scientific and technological intelligence based on multimodal data acquisition: the specific method for screening and collecting basic intelligence within the time frame is as follows:

[0018] When the intelligence timeliness index TI ≤ intelligence timeliness index threshold F, it means that the basic intelligence is within its time limit.

[0019] When the intelligence timeliness index TI > the intelligence timeliness index threshold F, it means that the basic intelligence has exceeded the timeliness limit.

[0020] In the preferred scheme of the above-mentioned method and system for collecting scientific and technological intelligence based on multimodal data acquisition: the specific steps for calculating the text index dex are as follows:

[0021] Text data includes several basic intelligence metrics: reliability score Sou, text integrity score Key, text content relevance score Con, text length len, and maximum text length len. max Minimum text length lenmin and the text release time Ti and the average text release time Ti ang ;

[0022] Calculate the text index dex of different basic information according to the text data. The specific formula is as follows:

[0023] ,

[0024] where α represents the weight coefficient of the source reliability score Sou, with a value of 0.1 < α < 0.4; β represents the weight coefficient of the text integrity score Key, with a value of 0.2 < β < 0.5; γ represents the weight coefficient of the content relevance score Con, with a value of 0.3 < γ < 0.4, and α + β + γ = 1; f represents the weight coefficient of, with a value of 0.1 < f < 1.

[0025] In the above preferred solution of the scientific and technological intelligence collection method and system based on multi-modal data: The specific steps for calculating the image index ma are as follows:

[0026] The image data includes the image reliability score cla, the image integrity score res, the image content relevance score tec, the minimum value of the resolution , the maximum value , the shooting time and the average shooting time ;

[0027] Calculate the image index ma of different basic information according to the image data. The specific formula is as follows:

[0028] ,

[0029] where x1 represents the weight coefficient of, with a value of 0.2 < x1 < 0.5, x2 represents the weight coefficient of, with a value of 0.1 < x2 < 0.4, x3 represents the image content relevance score of, with a value of 0.3 < x3 < 0.4, and x1 + x2 + x3 = 1.

[0030] In the above preferred solution of the scientific and technological intelligence collection method and system based on multi-modal data: The specific steps for calculating the average shooting time are as follows:

[0031] The image data also includes the number of images ;

[0032] According to the shooting time and the number of images , calculate the average shooting time , and the specific formula is as follows:

[0033] ,

[0034] Among them, represents the shooting time of the j-th image, j represents the ordinal number of the image, and the value range is .

[0035] In the above preferred solution of the scientific and technological intelligence collection method and system based on multi-modal data acquisition: The specific steps for calculating the audio index tex are as follows:

[0036] The audio data includes the sound quality score sou, the audio integrity score cty, and the audio technical relevance score nce of several basic information belonging to the audio data;

[0037] According to the audio data, calculate the audio index tex of different basic information, and the specific formula is as follows:

[0038] .

[0039] In the above preferred solution of the scientific and technological intelligence collection method and system based on multi-modal data acquisition: The specific steps for calculating the video index vid are as follows:

[0040] The video data includes the video reliability score vis, the video integrity score yj, and the video content relevance score toi of several basic information belonging to the video data;

[0041] According to the video data, calculate the video index vid of different basic information, and the specific formula is as follows:

[0042] ,

[0043] Among them, w1 represents the weight coefficient of the video reliability score , and the value range is 0.1 < w1 < 0.4; w2 represents the weight coefficient of the video integrity score , and the value range is 0.2 < w2 < 0.5; w3 represents the weight coefficient of the video content relevance score , and the value range is 0.3 < w3 < 0.4, and w1 + w2 + w3 = 1.

[0044] In the above preferred solution of the scientific and technological intelligence collection method and system based on multi-modal data acquisition: The method for screening basic information that meets the standards is as follows:

[0045] The set of basic information thresholds includes the text index threshold JV, the image index threshold GV, the audio index threshold TJ, and the video index threshold VJ;

[0046] When the text index dex < the text index threshold JV, it means that the basic information does not meet the text data collection standard, and the text data is re-collected and recalculated; when the text index dex ≥ the text index threshold JV, it means that the basic information meets the text data collection standard, and the basic information is retained.

[0047] When the image index ma < the image index threshold GV, it means that the basic information does not meet the image data acquisition standard, and the image data is reacquired and recalculated; when the image index ma ≥ the image index threshold GV, it means that the basic information meets the image data acquisition standard, and the basic information is retained.

[0048] When the audio index tex < audio index threshold TJ, it means that the basic information does not meet the audio data acquisition standard, and the audio data is re-acquired and recalculated; when the audio index tex ≥ audio index threshold TJ, it means that the basic information meets the audio data acquisition standard, and the basic information is retained.

[0049] When the video index vid < the video index threshold VJ, it means that the basic information does not meet the video data collection standard, and video data is re-collected and recalculated; when the video index vid ≥ the video index threshold VJ, it means that the basic information meets the video data collection standard, and the basic information is retained.

[0050] This invention also discloses a method and system for collecting scientific and technological intelligence based on multimodal data acquisition, including:

[0051] The data acquisition module is used to collect basic intelligence data, which includes several basic intelligence pieces.

[0052] The data analysis module is used to calculate the intelligence timeliness index (TI) of several basic intelligences based on basic intelligence data; set the intelligence timeliness index threshold (F); and filter and collect basic intelligence within the time limit based on the comparison result between the intelligence timeliness index (TI) and the intelligence timeliness index threshold (F), and classify the filtered basic intelligence into text data, image data, audio data and video data.

[0053] The data processing module is used to calculate the text index (dex) based on text data, and the image index (ma) and average shooting time based on image data. The audio index tex is calculated based on the audio data, and the video index vid is calculated based on the video data.

[0054] The data filtering module is used to set a set of basic intelligence thresholds. It compares the text index (dex), image index (ma), audio index (tex), and video index (vid) with the corresponding thresholds in the set of basic intelligence thresholds to filter basic intelligence that meets the criteria.

[0055] (III) Beneficial Effects

[0056] This invention provides a method and system for collecting scientific and technological intelligence based on multimodal data acquisition, which has the following beneficial effects:

[0057] (1) Collect basic intelligence data and gather original materials to provide rich resources to support intelligence collection, lay the foundation for accurate screening and in-depth analysis, and start a comprehensive and systematic intelligence mining process.

[0058] (2) Calculate and screen timely intelligence, calculate based on intelligence basis and compare with threshold, accurately screen out intelligence within time limit, classify into multimodal data, ensure the timeliness of intelligence, and divert the data for fine processing of each modality, thereby improving the targeting of processing.

[0059] (3) The reliability score sou, the text integrity score key, the text content relevance score con, the text length len, and the maximum text length len are used to determine the reliability score sou, the text integrity score key, the text content relevance score con, the text length len, and the maximum text length len. max Minimum text length len min Text publication time Ti and average text publication time Ti ang The calculated text index (dex) enables in-depth mining of scientific and technological intelligence within texts, avoiding the limitations of simple keyword searches and more accurately extracting valuable scientific and technological information, thus improving the accuracy and comprehensiveness of text intelligence analysis. This is achieved through image reliability score (cla), image integrity score (res), image content relevance score (tec), and minimum resolution. Maximum value Shooting time and average shooting time Calculating the image index *ma* for different basic intelligence greatly improves the accuracy and efficiency of identification, enabling a more comprehensive extraction of scientific and technological intelligence from image data; utilizing sound quality scoring *sou* and audio integrity scoring... and audio technology relevance score The resulting audio index tex can uncover deeper levels of scientific and technological information in audio, rather than being limited to simple basic parameter measurements. This improves the ability to extract valuable scientific and technological intelligence from audio data. By comprehensively considering video reliability score vis, video integrity score yj, and video content relevance score toi, a comprehensive quantitative video index vid is obtained. This avoids the problems of only crudely extracting parts of the video for analysis and processing audio and video separately. It can comprehensively and accurately extract scientific and technological intelligence from the video, thus improving the quality of video data intelligence analysis.

[0060] (4) By comparing the text index dex with the text index threshold JV, standard-compliant text data can be selected, avoiding low-quality text from interfering with subsequent analysis and improving processing accuracy and efficiency; by comparing the image index ma with the image index threshold GV, standard-compliant image information can be identified, and poor-quality or irrelevant images can be excluded, ensuring the reliability and effectiveness of image analysis; by comparing the audio index tex with the audio index threshold TJ, qualified audio data can be selected, avoiding analysis errors caused by poor audio quality and improving the accuracy of audio information processing; by comparing the video index vid with the video index threshold VJ, standard-compliant video data can be selected, avoiding low-quality videos from affecting analysis results and ensuring the reliability of video information processing. Attached Figure Description

[0061] Figure 1 This is a flowchart illustrating the scientific and technological intelligence collection method based on multimodal data acquisition according to the present invention. Detailed Implementation

[0062] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0063] Please see Figure 1 This invention provides a method and system for collecting scientific and technological intelligence based on multimodal data acquisition, including:

[0064] Step 1: Collect basic intelligence data, which includes several basic intelligence pieces.

[0065] Step 2: Based on the basic intelligence data, calculate the intelligence timeliness index TI for several basic intelligence items; set the intelligence timeliness index threshold F; based on the comparison between the intelligence timeliness index TI and the intelligence timeliness index threshold F, filter and collect basic intelligence within the time limit, and classify the filtered basic intelligence into text data, image data, audio data, and video data.

[0066] Step 201, the specific steps for calculating the Intelligence Timeliness Index (TI) are as follows:

[0067] The basic intelligence data includes intelligence release duration t, decay rate ki, expiration threshold vt, and the longest effective period Tmax.

[0068] It should be noted that the intelligence release duration t can usually be calculated by subtracting the intelligence release time from the current time.

[0069] The decay rate ki is obtained through the analysis of historical intelligence data, collecting multiple sets of intelligence value scores and corresponding time data. tn represents different points in time. Value scores representing different points in time can be obtained using professional intelligence analysis tools, such as Yuanting Technology's Defense Intelligence Classification Management System, K2 Visual Analysis and Judgment Software, and Huaxun Intelligence Analysis and Judgment System. These scores are then used to evaluate the value of each set of data. Take the natural logarithm Using linear regression and based on the principle of least squares, a straight line is fitted. For linear regression, the calculation formula is:

[0070] ,

[0071] in, This represents the nth time point. This represents the value of the intelligence at the nth point in time.

[0072] In this context , , , .

[0073] The obsolescence threshold vt is defined as follows: when the value of intelligence decreases to a certain percentage of its original value, the intelligence is considered to be of low value and can be identified as obsolete. For example, if the value of intelligence decreases to 30% of its original value, the intelligence is considered to be of low value, and the obsolescence threshold vt is set to 30%.

[0074] The longest effective period Tmax is calculated by analyzing intelligence and determining the average effective period supported by various types of intelligence after its release.

[0075] The intelligence timeliness index is calculated based on the intelligence release duration t, decay rate ki, expiration threshold vt, and longest effective period Tmax. The specific formula used is as follows:

[0076] .

[0077] It should be noted that the attenuation part The exponential function represents the rate at which the value of intelligence decays over time. (Threshold part) This part reflects the extent to which intelligence is nearing obsolescence, when near When the value approaches 0, it indicates that the intelligence is nearing its end. Multiplying the two parts above yields TI. When TI is less than or equal to the intelligence timeliness index threshold F, the intelligence is within its time limit; when TI is greater than F, the intelligence has exceeded its time limit. This calculation allows for a quantitative assessment of the timeliness of intelligence.

[0078] By introducing parameters such as intelligence release duration t, decay rate ki, expiration threshold vt, and longest effective period Tmax, the timeliness of intelligence can be comprehensively and scientifically quantified. This partially reflects the decay of intelligence value over time, while also considering the effective period of the intelligence. By comparing the calculated TI with the threshold F, it is possible to accurately determine whether the intelligence is within its effective period, avoiding the subjectivity and uncertainty of manual judgment. This helps to efficiently screen out intelligence with timeliness value, improve the efficiency and accuracy of intelligence processing, and ensure the quality of the intelligence used.

[0079] Step 202, the specific method for filtering and collecting basic intelligence within the time limit is as follows:

[0080] When the intelligence timeliness index TI ≤ intelligence timeliness index threshold F, the intelligence is within its time limit.

[0081] When the intelligence timeliness index TI > the intelligence timeliness index threshold F, the intelligence has exceeded the timeliness limit.

[0082] Step 3: Calculate the text index (dex) based on the text data, and calculate the image index (ma) and average shooting time based on the image data. The audio index tex is calculated based on the audio data, and the video index vid is calculated based on the video data.

[0083] Step 301: The specific steps for calculating the text index dex are as follows:

[0084] Text data includes several basic intelligence metrics: reliability score Sou, text integrity score Key, text content relevance score Con, text length len, and maximum text length len. max Minimum text length len min Text publication time Ti and average text publication time Ti ang .

[0085] It should be noted that the reliability score Sou for basic intelligence is calculated by counting the number of words in the text that belong to the keyword set, and then dividing that number by the total vocabulary size of the text to obtain the percentage of keyword occurrences. Assuming the total vocabulary size of the text is N and the number of keyword occurrences is ni, then the keyword occurrence percentage P = (ni / N) × 100. The percentage P is then mapped to a score of 0-100. For example, if P = 30%, then sou = 30.

[0086] The text integrity score key is calculated by dividing the number of words in the technical principle description section of the text by the total number of words in the text to obtain the proportion of technical principles. Assuming the number of words in the technical principle section is L1 and the total number of words is L, then the proportion of technical principles is P1 = (L1 / L) × 100. Similarly, the proportion of experimental data is calculated by counting the number of words in the experimental data description section (L2), and the proportion of experimental data is P2 = (L2 / L) × 100. The proportion of application scenarios is calculated by counting the number of words in the application scenario description section (L3), and the proportion of application scenarios is P3 = (L3 / L) × 100. The sum of these three proportions is the text integrity score key = P1 + P2 + P3.

[0087] The text content relevance score Con uses a bag-of-words model to convert the text to be evaluated and the target topic text into vectors. The bag-of-words model simply counts the frequency of each word in the text, forming a vector. The word embedding model, on the other hand, maps words to a low-dimensional vector space, capturing their semantic information. Assuming the text vector to be evaluated is... =(a1,a2,...,an), where the target topic text vector is... =(b1,b2,...,bn), then the cosine similarity The formula is:

[0088] ,

[0089] The cosine similarity value is mapped to a score from 0 to 100. For example, if the cosine similarity is 0.7, then con=70.

[0090] The length of a text (len) can usually be obtained by simply counting the number of characters, words, or sentences in the text. The maximum text length (len) is... max and the minimum text length len min The method involves calculating the text length len of all basic information, and then obtaining the maximum and minimum values ​​of the text length len as the maximum text length len. max and the minimum text length len min .

[0091] The publication time Ti of a text can usually be obtained from the text's original data. For example, in a web article, there may be a publication time tag; in a document file, there may be a record of the creation time or modification time.

[0092] Average Time to Publish Text (Ti) ang First, we need to obtain the publication time Ti of all texts, and then calculate the average of these publication times to get the average publication time Ti of the texts. ang。

[0093] The specific formula for calculating the text index (dex) of different basic intelligence based on text data is as follows:

[0094] ,

[0095] Where α represents the weighting coefficient of the source reliability score Sou, with a value of 0.1 < α < 0.4; β represents the weighting coefficient of the text integrity score Key, with a value of 0.2 < β < 0.5; γ represents the weighting coefficient of the content relevance score Con, with a value of 0.3 < γ < 0.4, and α + β + γ = 1; f represents The weighting coefficient is set to 0.1. <f<1。

[0096] It should be noted that, This section takes into account the three main characteristics of the text. This part takes into account the impact of text length. This part takes into account the impact of the text's publication time. The final text index is obtained by multiplying these three parts, which comprehensively considers factors such as the reliability of the text's source, keyword density, content relevance, text length, and publication time.

[0097] The formula comprehensively considers multiple factors, combining source reliability, keyword density, content relevance, text length, and publication time to fully evaluate text value. It avoids the one-sidedness of single-factor evaluation, and the weighting is reasonably set. Different weight coefficients are used to balance the importance of each factor, which can be flexibly adjusted according to actual needs to ensure that key factors have a reasonable impact on the results. The text length is normalized to ensure that texts of different lengths are treated fairly in the evaluation. The handling of publication time can also reflect the impact of the timeliness of the text on its value, which helps to screen out more valuable and timely texts.

[0098] Step 302: The specific steps for calculating the image exponent ma are as follows:

[0099] Image data includes several basic information components: image reliability score (CLA), image integrity score (RES), image content relevance score (TEC), and minimum resolution. Maximum value Shooting time and average shooting time .

[0100] It should be noted that the image reliability score (cla) categorizes and scores image sources, with officially released images receiving a score of 0.8-1.0, and user-uploaded images that have not been officially verified receiving a score of 0.3-0.5. The image integrity score (res) uses Adobe Photoshop to extract features from images, identify key image elements, and count the number of key image elements appearing. For example, in military base intelligence images, key image elements include military facilities, weapons and equipment, defensive structures, personnel and the number of active soldiers, as well as the base layout and infrastructure building structures. By identifying and counting these key image elements, important data support can be provided for intelligence analysis, helping to make accurate military assessments and decisions. The percentage of key image elements appearing is then calculated, and this percentage is the image integrity score (res).

[0101] Image content relevance scoring is performed using Adobe Photoshop (TEC). Image content analysis techniques are used to extract feature vectors from the image. These feature vectors are then compared to feature vectors from images related to the target theme. The relevance score (TEC) is determined by calculating the similarity. For example, first, a target theme, such as "beach scenery," is identified, and high-quality reference images related to it are collected. Next, the image to be evaluated is opened in Adobe Photoshop; if the image is not in a digital format, it can be converted. Adobe Photoshop's built-in image content analysis functions are used to extract features such as color histogram, texture, and shape, and these are combined into feature vectors. Simultaneously, feature vectors are extracted from the reference images using the same method. Then, an appropriate similarity metric, such as cosine similarity, is selected. The feature vectors of the image to be evaluated are compared one by one with the feature vectors of the reference images to calculate the similarity. Finally, scoring criteria are set; for example, a similarity between 0.8 and 1.0 corresponds to a relevance score of 90-100 points, and a similarity between 0.6 and 0.8 corresponds to a score of 70-90 points. Based on the calculated similarity and the set criteria, the image's relevance score (TEC) is determined, completing the assessment of image content relevance.

[0102] Minimum resolution Maximum value Record the minimum image resolution using image viewing software such as IrfanView. and maximum value .

[0103] Calculate the average shooting time The specific steps are as follows:

[0104] Image data also includes the number of images. .

[0105] Based on shooting time and the number of images , calculate the average shooting time , and the specific formula is as follows:

[0106] ,

[0107] Among them, represents the time of the j-th image, and j represents the ordinal number of the image, with a value range of .

[0108] Calculate the image index ma of different basic information based on the image data, and the specific formula is as follows:

[0109] Among them, x1 represents the weight coefficient of , with a value range of 0.2 < x1 < 0.5, x2 represents the weight coefficient of , with a value range of 0.1 < x2 < 0.4, x3 represents the image content relevance score the weight coefficient of , with a value range of 0.3 < x3 < 0.4, and x1 + x2 + x3 = 1.

[0110] It should be noted that this part takes into account the influence of the image reliability score and resolution, this part takes into account the influence of the image integrity score and shooting time, this part takes into account the relationship between the shooting time and the average shooting time, this part takes into account the influence of the image content relevance score. By adding these four parts, the image index is finally obtained, comprehensively considering factors such as image reliability, integrity, content relevance, resolution, and shooting time.

[0111] This formula combines multiple factors such as image reliability, integrity, content relevance, resolution, and shooting time, avoids the one-sidedness of single-factor evaluation, can comprehensively reflect the value of the image, balances the importance of each factor through different weight coefficients, can be flexibly adjusted according to actual needs, ensures that key factors have a reasonable impact on the results, and scientifically processes the shooting time and resolution. It takes into account the relationship between the shooting time and the average shooting time, as well as the normalization processing of the resolution, can more accurately evaluate the value of the image in terms of time and resolution, and helps to screen out high-quality images that meet the requirements.

[0112] Step 303: The specific steps for calculating the audio index tex are as follows:

[0113] The audio data includes the sound quality score sou, audio integrity score cty, and audio technology relevance score nce of several basic information belonging to the audio data.

[0114] It should be noted that the sound quality score (sou) was obtained using Adobe Audition audio processing software; the audio integrity score (cty) was obtained using Cool Edit Pro audio integrity processing software; and the audio technology relevance score (nce) was obtained using IBM Watson Speech to Text audio processing software, as follows: In the software interface, select the audio file to be analyzed, and preprocess the audio as needed. Then start the analysis. The software will perform speech recognition and transcription. After transcription, it will begin technology relevance analysis. On the one hand, it searches for audio technology-related keywords in the transcribed text, such as "audio encoding," and on the other hand, it performs semantic analysis to determine the depth and breadth of the text in the audio technology field. Then, it calculates the score, assigning weights to the identified keywords according to a preset keyword weight table. Simultaneously, it scores the depth of the text in the audio technology field based on the semantic analysis results. Finally, it combines the keyword weight score and the semantic depth score to obtain the audio technology relevance score (nce).

[0115] The specific formula for calculating the audio index tex for different basic information based on audio data is as follows:

[0116] .

[0117] It should be noted that, This means that the sound quality score, integrity score, and technical relevance score are added together, and then divided by 3 to obtain the average of the three audio files based on the sum of the three scores. The formula obtains the proportional values ​​related to the three scores. The final result of the formula is the sum of the two parts mentioned above to obtain the audio index tex. The first part is the average of the sum of the three scores, and the second part is the proportional value related to the sum of the products of the three scores, which comprehensively considers the audio's sound quality, integrity, and technical relevance.

[0118] The formula combines audio quality scores, integrity scores, and technical relevance scores, avoiding the one-sidedness of relying on a single factor to evaluate audio. It can comprehensively reflect the quality of audio. By summing and averaging, it can quantitatively compare the overall performance of different audio in these three indicators, which helps to screen out audio with better overall performance. The latter part of the formula further refines the evaluation mechanism by the proportional relationship between the product and sum of the three scores, taking into account the interrelationship between the scores, and can more accurately reflect the actual value of audio in multiple dimensions.

[0119] Step 304: The specific steps for calculating the video index vid are as follows:

[0120] The video data includes the video reliability score vis, the video integrity score yj, and the video content relevance score toi, which belong to several basic information of the video data.

[0121] It should be noted that the video reliability score vis can be obtained by existing video detection tools, such as Tencent Cloud Media Quality Inspection Software, etc., which detect whether the video covers 13 detection types such as screen freeze, black borders, mosaics, noise, etc., and provide an overall video quality detection score. By summing up the detections and scores of various aspects of the video, it is used as the video reliability score vis.

[0122] The video integrity score can be obtained through video parsing tools such as Yuanchuangbao, Doukuai Short Video Parser, etc. By detecting and analyzing non-original elements, it is judged whether there are problems affecting integrity such as content missing or being replaced in the video. A video without missing content is recorded as 100 points, and a video with suspected missing content is recorded as 50 points;

[0123] The video content relevance score toi is obtained through the IBM Watson Media video content analysis software. According to the target theme, the extracted keywords and the identified content are matched with the target theme. In terms of keyword matching, the frequency and importance of keywords related to the target theme appearing in the video are calculated. For example, for a theme about "the application of artificial intelligence in healthcare", the occurrences of keywords such as "medical image diagnosis" and "machine learning algorithms" in the video are counted. In terms of semantic analysis, by comprehensively understanding the semantics of the video content, the relevance between the video content and the target theme is judged. For example, it is analyzed whether the video content is discussing how artificial intelligence helps doctors diagnose diseases and other related content. According to the results of keyword matching and semantic analysis, the video content relevance score toi is calculated comprehensively.

[0124] The video index vid of different basic information is calculated based on the video data, and the specific formula is as follows:

[0125] ,

[0126] where, w1 represents the weight coefficient of the video reliability score with a value range of 0.1 < w1 < 0.4; w2 represents the weight coefficient of the video integrity score with a value range of 0.2 < w2 < 0.5; w3 represents the weight coefficient of the video content relevance score with a value range of 0.3 < w3 < 0.4, and w1 + w2 + w3 = 1.

[0127] It should be noted that the formula comprehensively evaluates the quality of the video by performing weighted summation on these three aspects of reliability, integrity, and relevance.

[0128] The reliability score (vis) multiplied by its weighting coefficient (w1) reflects the contribution of the video source's reliability to the video index. The completeness score (yj) multiplied by its weighting coefficient (w2) reflects the impact of the degree to which the video contains key information on the video index. The relevance score (toi) multiplied by its weighting coefficient (w3) represents the effect of the video content's relevance to the target topic on the video index. Finally, these three parts are added together to obtain the video index (vid). The higher the value, the higher the overall quality of the video.

[0129] This formula comprehensively and scientifically evaluates video quality by considering video reliability score (vis), integrity score (yj), and content relevance score (toi) and setting reasonable weighting coefficients. It avoids the one-sidedness of single-factor evaluation and limits the range of values ​​for the weighting coefficients to ensure that the proportion of each factor in the evaluation is scientific, making the evaluation results more accurate and reliable. This quantitative calculation method can easily and quickly filter and sort a large number of videos, and can efficiently identify high-quality videos in the information processing process, improving work efficiency and the accuracy of information processing.

[0130] Step 4: Set a set of basic intelligence thresholds. Compare the text index (dex), image index (ma), audio index (tex), and video index (vid) with the corresponding thresholds in the set of basic intelligence thresholds to filter out basic intelligence that meets the criteria.

[0131] The method for filtering basic information that meets the criteria is as follows:

[0132] The Text Index Threshold (JV) is established by collecting text data samples, such as official texts or documents, covering different types, fields, and quality levels. Examples include academic papers, news reports, and technical documents. For each text sample, its text index (dex) is calculated. Then, the dex values ​​of all text samples are aggregated, and the average of these dex values ​​is calculated. This average value serves as the standard for the Text Index Threshold (JV).

[0133] The Image Index Threshold (GV) is determined by collecting official image data samples that cover different types, fields, and quality levels, such as scientific instruments, technological products, and experimental scenarios. These samples should include different types of images, such as landscapes, portraits, product photos, and works of art, with varying image quality. For each image sample, its image index (ma) is calculated. The image index (ma) values ​​of all image samples are aggregated, and the average of these image index (ma) values ​​is calculated. This average value serves as the standard for the Image Index Threshold (GV).

[0134] The Audio Index Threshold TJ is determined by collecting official audio data samples that cover different types, fields, and quality levels of audio, such as live sound from a technology product launch, sound from the operation of scientific instruments, and recordings of technology lectures. These samples include music, speech, and environmental sound effects, and vary in quality. For each audio sample, its audio index tex is calculated. The audio index tex values ​​of all audio samples are then aggregated, and the average of these audio index tex values ​​is calculated. This average value serves as the standard for the Audio Index Threshold TJ.

[0135] The Video Index Threshold (VJ) is determined by collecting official video data samples that cover different types, fields, and quality levels of videos, such as movie clips, documentaries, advertisements, and home videos, with varying video quality. For each video sample, a Video Index (vid) is calculated. The average of the Video Index (vid) values ​​of all video samples is then calculated and used as the standard for the Video Index Threshold (VJ).

[0136] When the text index dex < the text index threshold JV, it means that the basic information does not meet the text data collection standard, and the text data is re-collected and recalculated; when the text index dex ≥ the text index threshold JV, it means that the basic information meets the text data collection standard, and the basic information is retained.

[0137] When the image index ma < the image index threshold GV, it means that the basic information does not meet the image data acquisition standard, and the image data is re-acquired and recalculated; when the image index ma ≥ the image index threshold GV, it means that the basic information meets the image data acquisition standard, and the basic information is retained.

[0138] When the audio index tex < audio index threshold TJ, it means that the basic information does not meet the audio data acquisition standard, and the audio data is re-acquired and recalculated; when the audio index tex ≥ audio index threshold TJ, it means that the basic information meets the audio data acquisition standard, and the basic information is retained.

[0139] When the video index vid < the video index threshold VJ, it means that the basic information does not meet the video data collection standard, and video data is re-collected and recalculated; when the video index vid ≥ the video index threshold VJ, it means that the basic information meets the video data collection standard, and the basic information is retained.

[0140] It should be noted that the entire process is a step-by-step screening process. By comparing the indices of text, images, audio, and video with thresholds, the correct data is selected. This method breaks down complex multimodal data processing into multiple single-modal data processing steps, making it easier to operate and manage.

[0141] By comparing data with thresholds, poor-quality data can be effectively eliminated. This modality-based, step-by-step filtering method helps optimize the data processing workflow. It enables data processors to process data of different modalities in a targeted manner, improving processing efficiency. At the same time, by setting thresholds for filtering, the subjectivity of human judgment is reduced, making the data filtering process more objective and scientific, and ensuring data consistency. In multimodal data processing, this filtering method helps ensure consistency between different modalities. Only when each modality of data passes its own filtering criteria will it proceed to the next step of processing, avoiding contradictions and mismatches between different modalities, thereby ensuring data consistency and usability.

[0142] On the other hand, the present invention also discloses a method and system for collecting scientific and technological intelligence based on multimodal data acquisition, including:

[0143] The data acquisition module is used to collect basic intelligence data, which includes several basic intelligence pieces.

[0144] The data analysis module is used to calculate the intelligence timeliness index (TI) of several basic intelligences based on basic intelligence data; set the intelligence timeliness index threshold (F); and filter and collect basic intelligence within the time limit based on the comparison between the intelligence timeliness index (TI) and the intelligence timeliness index threshold (F), and classify the filtered basic intelligence into text data, image data, audio data, and video data.

[0145] The data processing module is used to calculate the text index (dex) based on text data, and the image index (ma) and average shooting time based on image data. The audio index tex is calculated based on the audio data, and the video index vid is calculated based on the video data.

[0146] The data filtering module is used to set a set of basic intelligence thresholds. It compares the text index (dex), image index (ma), audio index (tex), and video index (vid) with the corresponding thresholds in the set of basic intelligence thresholds to filter basic intelligence that meets the criteria.

[0147] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.

[0148] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0149] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A method for collecting technology intelligence based on multi-modal data acquisition, characterized in that: Comprise: Collecting intelligence basic data, the intelligence basic data comprising several basic intelligence; According to the intelligence basic data, the intelligence time limit index TI of several basic intelligence is calculated; Set the intelligence time limit index threshold F, according to the comparison result of intelligence time limit index TI and intelligence time limit index threshold F, the basic intelligence in time limit is screened and collected, and the screened basic intelligence is classified to form text data, image data, audio data and video data;The specific steps of calculating intelligence time limit index TI are: The intelligence basic data includes intelligence release time t, decay rate ki, over time threshold vt and effective maximum period Tmax; According to the information release time length t, the decay rate ki, the out-of-date threshold vt and the effective maximum period Tmax, the information timeliness index is calculated The specific formula is as follows: calculating a text index dex from the text data, an image index ma from the image data and an average shooting time tim avg , an audio index tex from the audio data and a video index vid from the video data; The specific steps of calculating text index dex are: The text data includes a reliability score Sou, a text integrity score Key, a text content relevance score Con, a text length len, a maximum text length len max , a minimum text length len min , a text publishing time Ti, and a text average publishing time Ti ang ; According to the text data, the text index dex of different basic intelligence is calculated, and the specific formula is as follows: Wherein, a represents the weight coefficient of the source reliability score Sou, and the value is 0.1 < a < 0.4; β represents the weight coefficient of the text integrity score Key, and the value is 0.2 < β < 0.5; γ represents the weight coefficient of the content correlation score Con, and the value is 0.3 < γ < 0.4, and a + β + γ = 1; f represents the weight coefficient of the content correlation score Con, and the value is 0.1 < f < 1. the weight coefficient of the content correlation score Con, and the value is 0.1 < f < 1. The specific steps of calculating image index ma are: The image data comprises an image reliability score cla, an image integrity score res, an image content relevance score tec, a minimum value re min , a maximum value re max , a shooting time tim and an average shooting time tin avg ; According to the image data, the image index ma of different basic intelligence is calculated, and the specific formula is as follows: Among them, x1 represents the weight coefficient, with a value range of 0.2 < x1 < 0.

5. x2 represents the weight coefficient of tim, with a value range of 0.1 < x2 < 0.

4. x3 represents the weight coefficient of tec, with a value range of 0.3 < x3 < 0.4, and x1 + x2 + x3 = 1; The specific steps of calculating audio index tex are: Audio data includes the sound quality score sou, audio integrity score cty and audio technology relevance score nce of several basic intelligence belonging to audio data; According to the audio data, the audio index tex of different basic intelligence is calculated, and the specific formula is as follows: The specific steps of calculating video index vid are: Video data includes video reliability score vis, video integrity score yj and video content relevance score toi of several basic intelligence belonging to video data; According to the video data, the video index vid of different basic intelligence is calculated, and the specific formula is as follows: Wherein, w1 represents the weight coefficient of video reliability score vis, and the value is 0.1<w1<0.4;W2 represents the weight coefficient of video integrity score yj, and the value is 0.2<w2<0.5;W3 represents the weight coefficient of video content relevance score toi, and the value is 0.3<w3<0.4, and w1+w2+w3=1; Set the basic intelligence threshold set, compare text index dex, image index ma, audio index tex and video index vid with the corresponding threshold in basic intelligence threshold set respectively, and screen the basic intelligence meeting the standard, the method for screening the basic intelligence meeting the standard is: The basic intelligence threshold set includes text index threshold JV, image index threshold GV, audio index threshold TJ and video index threshold VJ; When text index dex is less than text index threshold JV, it indicates that this basic intelligence does not meet the text data collection standard, and the text data is collected again for calculation;When text index dex is greater than or equal to text index threshold JV, it indicates that this basic intelligence meets the text data collection standard, and the basic intelligence is retained; When image index ma is less than image index threshold GV, it indicates that this basic intelligence does not meet the image data collection standard, and the image data is collected again for calculation;When image index ma is greater than or equal to image index threshold GV, it indicates that this basic intelligence meets the image data collection standard, and the basic intelligence is retained; When the audio index tex is less than the audio index threshold TJ, it indicates that the basic information does not meet the audio data collection standard, and the audio data is re-collected for calculation; when the audio index tex is greater than or equal to the audio index threshold TJ, it indicates that the basic information meets the audio data collection standard, and the basic information is retained. When the video index vid is less than the video index threshold VJ, it indicates that the basic information does not meet the video data collection standard, and the video data is re-collected for calculation; when the video index vid is greater than or equal to the video index threshold VJ, it indicates that the basic information meets the video data collection standard, and the basic information is retained.

2. The technology intelligence gathering method based on multi-modal data acquisition according to claim 1, characterized in that: The specific method for screening and collecting the basic information within the time limit is: When the information time limit index TI is less than or equal to the information time limit index threshold F, it indicates that the basic information is within the time limit; When the information time limit index TI is greater than the information time limit index threshold F, it indicates that the basic information is beyond the time limit.

3. The multi-modal data acquisition based technology intelligence gathering method as claimed in claim 2, wherein: The average shooting time tim is calculated avg The specific steps are as follows: The image data further includes the image quantity A; The average shooting time tim is calculated from the shooting time tim and the number of images A avg The specific formula is as follows: where tim j denotes the time of taking the jth image, j denotes the ordinal number of the image, and takes the value [1, A].

4. A technology information collection system based on multi-modal data acquisition, which performs the technology information collection method based on multi-modal data acquisition according to any one of claims 1-3, characterized in that: The data collection module is configured to collect the information basic data, and the information basic data includes a plurality of basic information; The data analysis module is configured to calculate the information time limit index TI of the plurality of basic information according to the information basic data, set the information time limit index threshold F, and screen and collect the basic information within the time limit according to the comparison result of the information time limit index TI and the information time limit index threshold F, and classify the screened basic information to form the text data, the image data, the audio data and the video data; a data processing module configured to calculate a text index dex from the text data, an image index ma from the image data, and an average shooting time tim avg , an audio index tex from the audio data, and a video index vid from the video data; The data screening module is configured to set the basic information threshold set, compare the text index dex, the image index ma, the audio index tex and the video index vid with the corresponding threshold in the basic information threshold set respectively, and screen the basic information meeting the standard.

Citation Information

Patent Citations

  • Transportation information processing and sensing method based on multi-modal data fusion sensing

    CN117370932A