Scientific and technological data processing method and device
By acquiring and clustering technical terms from scientific and technological data, the system automatically identifies hot technical fields and quantifies changes, solving the problem of low efficiency in manual retrieval and achieving accurate prediction of scientific and technological trends.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING SCI & TECH PATENT OFFICE
- Filing Date
- 2025-12-16
- Publication Date
- 2026-04-24
AI Technical Summary
Current technologies for sensing technological trends rely on manual retrieval and analysis, which is inefficient and prone to missing important information, resulting in low accuracy.
By acquiring technology data from multiple time windows, extracting technical terms, using clustering to identify hot technology fields, and quantifying the degree of change in technology data, we can predict technology trends.
It has enabled automated and precise perception of technological trends, improving the accuracy and efficiency of technological trend perception.
Smart Images

Figure CN121919781A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method and apparatus for processing scientific and technological data. Background Technology
[0002] In this era of rapid technological advancement, the swift changes in technological dynamics and the emergence of massive amounts of information make timely and accurate perception of new trends, achievements, and challenges in the field of science and technology crucial. In related technologies, the perception of technological trends often relies on manual retrieval and analysis, which is inefficient and prone to overlooking important information. Therefore, the low accuracy and efficiency of technological trend perception remain pressing issues that need to be addressed. Summary of the Invention
[0003] This application provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for processing scientific and technological data, which can improve the accuracy and efficiency of scientific and technological trend perception.
[0004] The technical solution of this application embodiment is implemented as follows: This application provides a method for processing scientific and technological data, including: Acquire scientific and technological data for each of multiple time windows, the scientific and technological data including multiple technical terms; For each time window, multiple technical terms within the time window are clustered to obtain clusters, and the technical fields represented by the technical terms corresponding to the cluster centers of the clusters are taken as the hot technical fields of the time window. For each time window, determine the degree of change in the scientific and technological data of that time window compared to the historical scientific and technological data of historical time windows; Based on the degree of change in each of the time windows, the technological mutation events included in the scientific and technological data of the multiple time windows are determined; Based on the aforementioned technological upheaval events and the hot technological fields in each of the aforementioned time windows, technological trends are predicted.
[0005] This application embodiment also provides a scientific and technological data processing apparatus, including: The acquisition module is used to acquire scientific and technological data for each of the multiple time windows, the scientific and technological data including multiple technical terms; The clustering module is used to cluster multiple technical terms in each time window to obtain clusters, and to take the technical fields represented by the technical terms corresponding to the cluster centers of the clusters as the hot technical fields of the time window. The first determining module is used to determine, for each time window, the degree of change of the scientific and technological data of the time window compared with the historical scientific and technological data of the historical time window; The second determining module is used to determine the technological mutation events included in the scientific and technological data of the plurality of time windows based on the degree of change of each time window; The prediction module is used to predict technology trends based on the technological mutation events and the hot technology areas in each time window.
[0006] This application also provides an electronic device, including: Memory is used to store executable instructions for a computer; The processor, when executing computer-executable instructions stored in the memory, implements the method for processing scientific and technological data provided in the embodiments of this application.
[0007] This application also provides a computer-readable storage medium storing computer-executable instructions or computer programs, which, when executed by a processor, implement the scientific and technological data processing method provided in this application.
[0008] This application also provides a computer program product, including computer-executable instructions or a computer program, which, when executed by a processor, implements the scientific and technological data processing method provided in this application.
[0009] The embodiments of this application have the following beneficial effects: First, by acquiring scientific and technological data from multiple time windows and extracting technical terms, and automatically identifying hot technical fields in each time window through clustering, the efficiency of scientific and technological trend perception is improved, avoiding the omission of important information by humans. Second, by determining the degree of change of scientific and technological data in each time window compared to historical time windows, technological changes are quantified and technological mutation events are accurately identified, solving the problem that it is difficult for humans to accurately judge technological mutations. Finally, based on technological mutation events and hot technical fields, technological trends are predicted, integrating key information on technological dynamics (i.e., technological mutation events and hot technical fields), thus improving the accuracy of technological trend prediction. In summary, the embodiments of this application achieve automated and precise perception of scientific and technological trends, improving the accuracy and efficiency of scientific and technological trend perception. Attached Figure Description
[0010] Figure 1 This is a schematic diagram of the architecture of the scientific and technological data processing system provided in the embodiments of this application; Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application; Figure 3This is a first flowchart illustrating the method for processing scientific and technological data provided in the embodiments of this application; Figure 4 This is a second flowchart illustrating the method for processing scientific and technological data provided in the embodiments of this application; Figure 5 This is a schematic diagram of the third process of the scientific and technological data processing method provided in the embodiments of this application; Figure 6 This is a schematic diagram of the fourth process of the scientific and technological data processing method provided in the embodiments of this application; Figure 7 This is a schematic diagram of the process for generating second lexical features provided in an embodiment of this application; Figure 8 This is a flowchart illustrating the identification of hotspot technology areas provided in an embodiment of this application; Figure 9 This is a flowchart illustrating the mutation event of the marker technology provided in the embodiments of this application; Figure 10 This is a flowchart illustrating the process of predicting technological trends provided in an embodiment of this application.
[0011] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0014] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0015] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of a larger module or unit that includes the functionality of the module or unit.
[0016] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.
[0017] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0018] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0019] 1) In response to, used to indicate the conditions or states on which the operation performed depends. When the conditions or states on which it depends are met, one or more operations performed may be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations are performed.
[0020] 2) Tech data refers to unstructured / semi-structured raw data used to analyze technological dynamics. It covers core information sources in the technology field, such as: a) Academic: paper abstracts / full texts, conference reports; b) Industry: technology news, corporate white papers; c) Intellectual property: patent texts. Tech data is the source of information for technological development, and subsequent analysis relies on the semantic content of the tech data.
[0021] 3) Tech terms are the core semantic carriers of scientific and technological data. They are terms extracted from scientific and technological data that represent technical concepts or entities (such as multimodal input interfaces and long text processing optimization). They are the semantic core of scientific and technological dynamics. Content unrelated to the semantic core (such as press conference locations and author names) will be filtered out, and only terms directly related to technology will be retained.
[0022] 4) Large Language Model (LLM) is a deep learning-based natural language processing model, typically employing the Transformer architecture and trained on large-scale text corpora (such as books and web pages). Its core capabilities include understanding natural language semantics, generating context-appropriate text, and capturing semantic relationships within long contexts. This model supports various tasks such as dialogue interaction, text translation, content creation, and information retrieval, and can simulate human language understanding and generation capabilities, making it one of the key technologies in the field of natural language processing.
[0023] 5) Technological trends refer to time-series analysis of scientific and technological data. This involves dividing time windows, clustering technical terms to identify hot technology areas, quantifying the degree of change in scientific and technological data between the current and historical time windows to detect technological upheaval events, and integrating hot technology areas and technological upheaval events to predict the direction or trend of technological development through models. Its core is to use automated and quantitative methods to integrate the "conventional trends (i.e., hot technology areas)" and "mutational disturbances (i.e., technological upheaval events)" of technological development, providing objective and accurate judgments on future technological development for the dynamic perception of science and technology. Specifically, technological trends can be characterized by technological trend indicators. Therefore, technological trend prediction maps the convergence event signals of hot technology fields and technological mutation events into quantifiable technological trend indicators. For example, technological trend indicators include, but are not limited to: a) Technological importance I_pred: a continuous value, ranging from 0 to 1, representing the future importance of the technology field; b) Growth rate G_pred: a continuous value (e.g., 0.15), representing the growth rate of the importance of the technology field (e.g., a 15% increase); c) Technology field size S_pred: a continuous value, representing the scale of scientific and technological data in the technology field, such as the number of documents / patents; d) Mutation risk probability P_pred: ranging from 0 to 1, representing the probability of a technological mutation event occurring in the future; e) Confidence interval [low, high]: representing the confidence range of the predicted value of the technological trend indicator, such as the confidence interval of I_pred [0.88, 0.95]).
[0024] This application provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for processing scientific and technological data, which can improve the accuracy and efficiency of scientific and technological trend perception. The following is a detailed description of the embodiments of this application based on the above explanation of the terms and concepts used.
[0025] The following describes the scientific and technological data processing system provided in the embodiments of this application. See also Figure 1 , Figure 1This is a schematic diagram of the architecture of a scientific and technological data processing system provided in an embodiment of this application. To support an exemplary application, the scientific and technological data processing system 100 includes: a server 200, a network 300, a terminal 400, and a database 600. The terminal 400, server 200, and database 600 are connected via the network 300, which can be a wide area network (WAN), a local area network (LAN), or a combination of both, using wireless or wired links to achieve data transmission.
[0026] Here, terminal 400 responds to the processing instruction for scientific and technological data by sending a processing request to server 200; server 200 responds to the processing request by retrieving scientific and technological data from database 600 for each of multiple time windows, the scientific and technological data including multiple technical terms; for each time window, the multiple technical terms of the time window are clustered to obtain clusters, and the technical fields represented by the technical terms corresponding to the cluster centers of the clusters are taken as the hot technical fields of the time window; for each time window, the degree of change of the scientific and technological data of the time window compared with the historical scientific and technological data of historical time windows is determined; based on the degree of change of each time window, the technological mutation events included in the scientific and technological data of multiple time windows are determined; based on the technological mutation events and the hot technical fields of each time window, the technological trend is predicted; the predicted technological trend is returned to terminal 400; terminal 400 receives the technological trend sent by server 200 and displays the predicted technological trend.
[0027] The method for processing scientific and technological data provided in this application embodiment is implemented by an electronic device. For example, it can be implemented by a terminal alone, by a server alone, or by a terminal and a server working together. The electronic device implementing the method for processing scientific and technological data provided in this application embodiment can be various types of terminals or servers. The server (e.g., server 200) can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal (e.g., terminal 400) can be a laptop, tablet, desktop computer, smartphone, intelligent voice interaction device (e.g., smart speaker), smart home appliance (e.g., smart TV), smartwatch, vehicle terminal, wearable device, virtual reality (VR) device, aircraft, etc., but is not limited to these. The terminal and server can be connected directly or indirectly through wired or wireless communication, and this application embodiment does not impose any restrictions on this.
[0028] In some embodiments, the terminal or server can implement the scientific and technological data processing method provided in this application by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be native applications (APPs), i.e., programs that need to be installed in the operating system to run; or they can be applets that can be embedded in any APP, i.e., programs that only need to be downloaded to a browser environment to run. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin.
[0029] The following describes an electronic device that implements a method for processing scientific and technological data, as provided in an embodiment of this application. See also... Figure 2 , Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 500 provided in this embodiment can be a terminal or a server. Figure 2 As shown, electronic device 500 includes at least one processor 510, memory 550, at least one network interface 520, and user interface 530. The various components in electronic device 500 are coupled together via a bus system 540. It is understood that the bus system 540 is used to implement communication between these components. In addition to a data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 2 The general labeled all buses as Bus System 540.
[0030] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0031] User interface 530 includes one or more output devices 531 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0032] Memory 550 may be removable, non-removable, or a combination thereof. Memory 550 may include one or more storage devices physically located away from processor 510. Memory 550 may include volatile memory or non-volatile memory, or both. Non-volatile memory may be read-only memory (ROM), and volatile memory may be random access memory (RAM). The memory 550 described in this application embodiment is intended to include any suitable type of memory.
[0033] In some embodiments, memory 550 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, as illustrated below. Operating system 551 includes system programs for handling various basic system services and performing hardware-related tasks, such as a framework layer, core library layer, driver layer, etc., for implementing various basic services and handling hardware-based tasks; network communication module 552 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 520, exemplary network interfaces 520 including Bluetooth, Wireless Fidelity (Wi-Fi), and Universal Serial Bus (USB); presentation module 553 is used to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 531 associated with user interface 530 (e.g., a display screen, a speaker, etc.); input processing module 554 is used to detect and translate one or more user inputs or interactions from one or more input devices 532.
[0034] In some embodiments, the scientific data processing apparatus provided in this application can be implemented in software. Figure 2 A processing device 555 for scientific and technological data stored in memory 550 is shown. It may be software in the form of programs and plug-ins, including the following software modules: acquisition module 5551, clustering module 5552, first determination module 5553, second determination module 5554, and prediction module 5555. These modules are logically related and can therefore be arbitrarily combined or further divided according to the functions they implement. The functions of each module will be described below.
[0035] The following describes the method for processing scientific and technological data provided in the embodiments of this application. As mentioned above, the method for processing scientific and technological data provided in the embodiments of this application is implemented by an electronic device, such as a server or terminal alone, or a server and terminal working together. Therefore, the executing entity of each step will not be described again below. See Figure 3 , Figure 3 This is a first flowchart illustrating the method for processing scientific and technological data provided in this application embodiment. The method for processing scientific and technological data provided in this application embodiment includes: Step 101: Obtain the scientific and technological data for each time window across multiple time windows.
[0036] Among them, the scientific and technological data includes a number of technical terms.
[0037] For step 101, (1) a time window is a method of dividing a continuous time series into discrete time periods of fixed duration (such as weekly, monthly, or quarterly), with each time period called a time window. Example: If the window size of the time window is set to "monthly", then... =January 20XX February 20XX... =December 20XX, each time window corresponds to all the technology data for that month. (2) Technology data is unstructured / semi-structured raw data used to analyze technology dynamics, covering the core information sources in the technology field, such as a) academic: paper abstracts / full texts, conference reports; b) industry: technology news, corporate white papers; c) intellectual property: patent texts. Technology data is the source of information for technological development, and subsequent analysis depends on the content semantics of technology data. (3) Tech terms are the core semantic carriers of technology data. They are terms extracted from technology data that represent technical concepts or technical entities (such as multimodal input interfaces, long text processing optimization). They are the semantic core of technology dynamics. Content unrelated to the semantic core (such as press conference locations, author names) will be filtered out, and only terms directly related to technology will be retained. For example, technical terms can be automatically extracted from scientific and technological data through data preprocessing and Large Language Model (LLM) technology. For instance, the text (i.e., scientific and technological data) can be segmented using the LLM tokenization interface to preserve the integrity of compound technical terms (e.g., "multimodal input interface" is not split into "multimodal / state / input / interface"); then, prompt word engineering can guide the LLM to output structured entities (i.e., technical terms, such as multimodal input interface, long text processing optimization).
[0038] In some embodiments, before performing step 101 "acquiring scientific and technological data for each time window in multiple time windows", the following steps may also be performed: obtaining the first window size and window adjustment coefficient of the initial time window; obtaining the time change of the i-th time window in multiple time windows compared to the initial time window, and determining the second window size of the i-th time window based on the first window size, window adjustment coefficient and time change, and determining the i-th time window with the second window size; traversing i to obtain multiple time windows, the multiple time windows being continuous in time, the starting time point of the first time window being the starting time point of the initial time window, i being an integer greater than 0 and less than or equal to N, and N being the number of time windows.
[0039] Here, (1) the initial time window is the "baseline time range" of the first time window, which can include the initial start time (start_time) and the initial duration (i.e., the size of the first window). Both the start time and the initial duration can be preset. (2) The window adjustment coefficient can determine the rate at which the window size changes over time. For example, if the window adjustment coefficient is greater than 1, the time window will be larger than the initial time window. If the window adjustment coefficient is less than 1, the time window will be smaller than the initial time window. The window adjustment coefficient can be related to the technical field to which the scientific and technological data belongs. Each technical field has a corresponding window adjustment coefficient. For example, the window adjustment coefficient (e.g., 0.8) of the rapid iteration technical field (e.g., the large model field, the quantum computing field, the artificial intelligence field) is smaller than the window adjustment coefficient (e.g., 1.2) of the slow iteration field (e.g., the traditional materials field), so that the time window of the rapid iteration technical field is smaller than the time window of the slow iteration field. (3) The time change is the time offset of the i-th time window relative to the initial time window (i.e., the total duration from the initial start time to the start time of the i-th time window, which is also the sum of the total durations of the i-th time window and the first i-1 time windows). The time change can be preset. (4) The second window size is the window size of the i-th time window.
[0040] Specifically, first obtain the size of the first window of the initial time window. Window adjustment coefficient The number of time windows N and the time change of the i-th time window compared to the initial time window. Therefore, based on the first window size, the window adjustment factor, and the time change, the second window size of the i-th time window is determined using the following formula (1): Formula (1) in, Let be the size of the second window of the i-th time window.
[0041] Continuing, after obtaining the size of the second window of the i-th time window, the i-th time window with that second window size is determined, thus identifying the i-th time window. Therefore, by iterating through i (where i is an integer greater than 0 and less than or equal to N, incrementing from 1 with a step size of 1), N time windows can be obtained. It should be noted that the starting time of the 1st (i=1) time window is the starting time of the initial time window, that is, the time change of the 1st time window compared to the initial time window. Furthermore, the multiple time windows must be continuous in time, meaning they must be consecutive without gaps. Therefore: the start time of the i-th time window = the end time of the (i-1)-th time window (the start time of the first window = the initial start time start_time); the end time of the i-th time window = the start time of the i-th time window + ... .
[0042] By applying the above embodiments and dynamically adjusting the size of the time window, the pain points of fixed time windows—either missing rapidly changing details (such as in the AI field) or resulting in fragmented data (such as in traditional manufacturing)—are addressed. Furthermore, continuous, uninterrupted time coverage ensures no gaps in scientific and technological data are missing, guaranteeing the completeness of the analysis. Combined with domain-specific window adjustment coefficients, a balance is achieved between "detail capture" and "data efficiency," ultimately providing more accurate foundational data for subsequent hotspot identification and mutation detection, improving the accuracy and adaptability of dynamic scientific and technological perception.
[0043] Step 102: For each time window, cluster the multiple technical terms in the time window to obtain clusters, and take the technical technology represented by the technical terms corresponding to the cluster center of the cluster as the hot technical technology of the time window.
[0044] Step 102 involves automatically identifying structured hot technology areas from scattered technical terms within a time window. First, for multiple technical terms within each time window (e.g., multimodal input, long text optimization), clustering algorithms (e.g., peak density clustering, k-means clustering) are used to cluster these terms, forming clusters (e.g., semantically similar technical terms form a cluster). Next, the cluster center of each cluster (corresponding to the technical term that best represents the core semantics of the cluster, such as multimodal input) represents a technology area, thus determining the technology area represented by the cluster center. Finally, the technology area represented by the cluster center is designated as the hot technology area for that time window (e.g., the technology area corresponding to the technical term "multimodal input" is "multimodal large model"). This automatically identifies the hot technology areas for each time window without manual retrieval and analysis, solving the problems of scattered technical terms and difficulty in structurally identifying hot topics. For example, if the technical terms within the time window include "large model A", "multimodal input interface" and "XXX architecture", after clustering, "large model A" and "multimodal input interface" form a cluster (with "GPT-4" as the cluster center), corresponding to the hot technical field of "multimodal large model"; "XXX architecture" forms a separate cluster, corresponding to the hot technical field of "open source large model".
[0045] In some embodiments, see Figure 4 The process of "clustering multiple technical terms within a time window to obtain clusters" can be achieved by performing the following steps 201-203: Step 201: For each technical term within the time window, determine the importance of the technical term to the target technical description statement. The scientific and technological data within the time window includes the target technical description statement, which in turn includes technical terms. Step 202: For each technical term, extract its first lexical feature and weight it using importance as the weight value to obtain its second lexical feature. Step 203: Cluster the multiple second lexical features to obtain clusters.
[0046] Step 201 involves determining the semantic contribution of technical terms to the target technical description statement. The target technical description statement is the one selected from the scientific and technological data in step 101 that includes technical terms (e.g., "methodological innovation sentences" in paper abstracts, "technical detail sentences" in news articles). For example, importance calculation can be achieved using an attention mechanism, such as the attention mechanism of a Large Language Model (LLM). Each technical term in the target technical description statement is then given attention to obtain its importance, which is a weighted value between 0 and 1. The greater the contribution of a technical term to the core semantics of the target technical description statement, the higher its importance.
[0047] Step 202 combines the "semantic information" and "importance" of technical terms to form a more accurate feature representation, thereby making the features of important terms more prominent. Here, the following processing is performed for each technical term: First, the first lexical feature of the technical term is extracted. The first lexical feature is the semantic vector of the technical term. This is obtained by mapping the technical term, for example, through the Text Embedding interface of LLM. Then, the importance of the technical term is used as the weight value of the first lexical feature, and the first lexical feature is weighted to obtain the second lexical feature of the technical term, i.e., second lexical feature = weight value * first lexical feature.
[0048] In step 203, the weighted second lexical features are used instead of the "naked semantic features (i.e., unweighted first lexical features)" for clustering to obtain clusters. For example, density peak clustering algorithm can be used to cluster multiple second lexical features to obtain clusters. In this way, the weighted second lexical features better reflect the true value (i.e., importance) of technical terms. The clusters contain important and relevant technical terms, and the cluster centers are the hot technical fields, making the clustering results more consistent with the definition of "hot technical fields".
[0049] By applying the above embodiments, the pain point of "naked semantic feature clustering being easily interfered with by irrelevant words" is solved through the link of "importance differentiation - weighted enhancement - precise clustering": Step 201 uses importance to filter out "the most core technical terms in the description of the target technology"; Step 202 uses importance as the weight to weight the first lexical features of the technical terms, making the semantic influence of technical terms with high importance (such as those above the importance threshold) more prominent; Step 203 clusters based on the weighted second lexical features, ensuring that "semantically related and important terms" are clustered into precise clusters. Ultimately, the clustering results are closer to the real technology associations, providing a reliable "analysis unit" for subsequent hotspot identification and mutation detection, and significantly improving the accuracy of technology trend perception.
[0050] In some embodiments, before performing the step "clustering multiple technical terms within a time window to obtain clusters", the following steps may also be performed: identifying multiple technical entities from scientific and technological data and encoding the multiple technical entities to obtain technical entity features; extracting multiple technical description statements including at least one technical entity from the scientific and technological data and encoding each technical description statement to obtain statement features; determining the feature similarity between the statement features and technical entity features of each technical description statement; and selecting the technical description statement whose feature similarity satisfies the similarity condition among the multiple technical description statements as the target technical description statement.
[0051] Here, the technical description statements are "purified" before clustering, that is, the target technical description statements that can truly represent the core of the technology are selected to avoid irrelevant content from interfering with subsequent analysis. (1) Identify technical entities from scientific and technological data. Technical entities are the "anchors" of scientific and technological data, representing specific technical concepts. The technical entities are encoded into technical entity features. Technical entity features are semantic vectors. For example, technical entities can be encoded using the text embedding function of LLM to obtain technical entity features. In this way, the "semantics" of technical entities are transformed into a computable numerical representation (i.e., technical entity features). In practical applications, multiple technical entities can be encoded separately to obtain their respective encoded features, and then the average of multiple encoded features can be calculated to obtain technical entity features; or multiple technical entities can be encoded separately to obtain their respective encoded features, and then multiple encoded features can be concatenated to obtain technical entity features. (2) Extract multiple technical description statements containing at least one technical entity from the scientific and technological data. Each technical description statement contains at least one technical entity. Then encode the technical description statements into statement features. For example, use an LLM text vector model (e.g., bge-large-zh) to encode the technical description statements to obtain semantic vectors (i.e., statement features), thereby converting the "semantics" of the technical description statements into numerical representations. (3) Calculate the feature similarity (e.g., cosine similarity) between the statement features of each technical description statement and the technical entity features corresponding to that technical description statement. The higher the feature similarity, the closer the core semantics of the technical description statement is to the technical entity. (4) Select technical description statements whose feature similarity meets the similarity condition from multiple technical description statements, and use the technical description statements whose feature similarity meets the similarity condition as the target technical description statements. Among them, the feature similarity meets the similarity condition, which means: the feature similarity is greater than the similarity threshold (e.g., 0.7); and a specific number of feature similarities that are ranked first when the feature similarities are sorted in descending order.
[0052] Applying the above embodiments, the core of the "technical content purification" of scientific and technological data is achieved through "technical entity anchoring - sentence association filtering": first, technical entities are identified and encoded; then, technical description sentences containing technical entities are extracted, and the semantic similarity (i.e., feature similarity) between the technical description sentences and the technical entities is calculated; target technical description sentences with strong associations (feature similarity satisfying the similarity condition) are retained. This process filters out irrelevant content in the scientific and technological data, avoiding irrelevant words from interfering with subsequent importance calculations and clustering; it allows the analysis to focus only on target technical description sentences that "truly involve technical details," improving the accuracy of technical word importance and feature extraction, laying a high-purity data foundation for clustering and hotspot identification, and enhancing the accuracy of scientific and technological trend perception.
[0053] In some embodiments, the step "encoding each technical description statement to obtain statement features" can be achieved by performing the following steps: encoding each technical description statement using the text vector model of a large language model to obtain statement features; correspondingly, the step 201 "determining the importance of technical terms to the target technical description statement" can be achieved by performing the following steps: performing attention processing on technical terms using the attention processing model of a large language model to obtain the importance of technical terms to the target technical description statement; correspondingly, the step 202 "extracting the first lexical features of technical terms" can be achieved by performing the following steps: performing text vectorization processing on technical terms using a large language model to obtain the first lexical features of technical terms.
[0054] Here, when encoding sentence features, a text vector model (such as bge-large-zh) of a Large Language Model (LLM) can be used to encode each technical description sentence to obtain sentence features. Correspondingly, when calculating importance, an attention processing model of the LLM can be used to apply attention processing to technical terms, obtaining the importance of each technical term to the target technical description sentence. Similarly, when extracting the first lexical feature, the LLM (specifically, text embedding processing) can be used to perform text vectorization processing on the technical terms to obtain the first lexical feature of the technical terms. Thus, the LLM achieves sentence feature encoding, importance calculation, and first lexical feature extraction.
[0055] By applying the above embodiments and implementing feature encoding and importance calculation through Large Language Model (LLM), the core value lies in enhancing the depth and accuracy of semantic understanding: encoding sentence / lexical features using an LLM text vector model can more accurately capture deep semantic relationships; calculating lexical importance using an LLM attention mechanism can automatically distinguish between "core technical terms" and "irrelevant terms" in a sentence. This approach replaces traditional feature extraction, making the features of technical terms more closely aligned with real semantics, and the importance more reflective of actual contributions. This lays a more accurate data foundation for subsequent clustering and hotspot identification, ultimately improving the accuracy and reliability of technology trend perception.
[0056] Next, combine Figure 7 ( Figure 7 This is a schematic diagram of the process for generating second lexical features provided in the embodiments of this application (detailed explanation of weighted lexical feature vectors). The generation process of (i.e., the second lexical feature mentioned above) includes: (1) Preprocessing of scientific and technological data: Lightweight cleaning of the collected unstructured scientific and technological data (such as academic papers, scientific and technological news, etc.) and using the LLM Tokenization interface to perform sub-word segmentation of scientific and technological data to ensure the complete parsing of compound terms in the field of science and technology.
[0057] (2) Feature extraction driven by Large Language Model (LLM). That is, based on the reasoning ability of the large language model, weighted lexical feature vectors are generated. Specifically, it includes: a) Perform basic semantic feature extraction to identify technical entities such as technical terms, research institutions, and research methods in scientific and technological data. For example, ... Figure 7 As shown, prompt word engineering can guide LLM to extract technical entities from scientific and technological data. For example, the output format of technical entities is: {"Terminology": ["Large Model A", "Multimodal Input Interface", "XXX Architecture"], "Institution": ["Institution 1", "Institution 2"], "Method": ["Long Text Processing Optimization"]}.
[0058] b) such as Figure 7 As shown, firstly, based on the technical entities identified in step a, technical description statements containing technical entities are identified; then, based on the text vector model of LLM (such as bge-large-zh), sentence vectors (i.e., sentence features) of technical description statements are generated, and the cosine similarity between sentence vectors and technical entities is calculated according to the cosine similarity algorithm; finally, the core technical description statement (i.e., the target technical description statement mentioned above) is selected from multiple technical description statements by combining the cosine similarity.
[0059] c) such as Figure 7 As shown, the following processing is performed on each core technology description statement: First, based on the keyword list, technical terms are extracted from the core technology description statements. Then, using the text embedding function of LLM, each technical term in the core technology description statement is mapped to a high-dimensional vector space to obtain the word vector of each technical term (i.e., the first word feature mentioned above, including...). , ... Word vectors contain semantic information of technical terms, and the distance between word vectors reflects the semantic similarity between words. Then, the importance measure (i.e., the aforementioned importance) of each technical term in the core technical description statement is calculated using an LLM-based attention mechanism, namely, the word-level attention weight (including...). , ... The value typically ranges from 0 to 1. Therefore, a weighted vocabulary feature vector is obtained. As shown in formula (2): Formula (2) in, It is a lexical feature vector that comprehensively considers the semantic information and importance of each technical term in the core technology description statement. It can more accurately represent the core semantics in the core technology description statement and can be used for subsequent tasks such as hotspot identification, technology mutation detection, and trend prediction.
[0060] Next, combine Figure 8 ( Figure 8 This is a flowchart illustrating the process of identifying hotspot technologies according to embodiments of this application, detailing the implementation process of identifying hotspot technologies, including: To capture trending technology sectors in real time and dynamically, a sliding time window is first set up, and then the following processing is performed on the technology data for each time window: Using the above... Figure 7 The processing flow generates lexical feature vectors for technical terms in the scientific and technological data within that time window. Then, the density peak clustering algorithm is used to cluster multiple word feature vectors corresponding to the time window to obtain multiple technology clusters (i.e., the above clusters, including technology cluster C1, technology cluster, ..., technology cluster Cn). In this way, data points with significantly higher density than the surrounding data points (i.e., word feature vectors) and far away from other high-density data points are identified as the cluster centers of the technology clusters (i.e., the above cluster centers). The technology fields represented by the cluster centers are the hot technology fields.
[0061] Each technology cluster C includes a central term (i.e., the technical term corresponding to the cluster center). and a set of member words (i.e., technical words in the technical cluster other than those corresponding to the cluster center) { , ... The importance of technology cluster C I (C) Quantifying the importance of the central words through word vector paradigms ,Right now: ,in, It is the central word Word vectors.
[0062] Finally, all hot technology areas within the time window are obtained, and the following information is generated: the core keyword of each hot technology area (i.e., the core keyword of the technology cluster C corresponding to the hot technology area), and the importance of each hot technology area (i.e., the importance of the technology cluster C). ), relevant technical terms for each hot technology field (i.e., technical terms included in the technology cluster C corresponding to the hot technology field); and, all technical terms within the time window and the importance of each technical term (which can be determined by the distance between the technical term and the central word (such as the central word of the technology cluster to which the technical term belongs; the greater the distance, the lower the importance).
[0063] Step 103: For each time window, determine the degree of change in the technological data of the time window compared to the historical technological data of the historical time windows.
[0064] Step 103 involves quantifying the differences in scientific and technological data between each time window and historical time windows, providing crucial evidence for subsequent "mutation event detection" and "technology trend prediction." Specifically, for each time window... Calculate the scientific and technological data for this time window and compare it with historical time windows (such as historical time windows). (m is an integer greater than 0); such as the average of scientific and technological data from multiple historical time windows. The greater the degree of change, the more significant the difference between the scientific and technological data of the current time window and the historical scientific and technological data of previous time windows. For example, the degree of change can be quantified through semantic similarity (such as the cosine similarity between the features of the scientific and technological data of the current time window and the features of the historical scientific and technological data of previous time windows; the lower the cosine similarity, the greater the degree of change), frequency difference (KL divergence between the frequency distribution of technical terms in the current time window and the frequency distribution of technical terms in previous time windows; the higher the KL value, the greater the degree of change), and cluster overlap (the proportion of intersection between the clusters corresponding to the current time window and the clusters corresponding to the previous time window; the lower the proportion, the greater the degree of change), etc. There are no restrictions here. In this way, transforming "technological change" from a qualitative description to a quantitative indicator is a key bridge connecting "static window analysis" (step 102) and "dynamic trend judgment," ensuring that subsequent analysis is based on verifiable "degree of change" rather than subjective judgment.
[0065] In some embodiments, step 103, "determining the degree of change of the scientific and technological data of the time window compared with the historical scientific and technological data of the historical time window," can be achieved by performing the following steps: determining a first historical time window adjacent to the time window, and determining a second historical time window and a third historical time window; wherein, the second historical time window is K time windows away from the time window, and the third historical time window is located before the first historical time window and is K time windows away from the first historical time window, the historical time window includes the first historical time window, the second historical time window, and the third historical time window, and K is an integer greater than 1; determining a first similarity between the scientific and technological data of the time window and the historical scientific and technological data of the second historical time window, and determining a second similarity between the historical scientific and technological data of the first historical time window and the historical scientific and technological data of the third historical time window; subtracting the second similarity from the first similarity to obtain the degree of change.
[0066] Here, we first determine the historical time window; specifically, we determine the time window... (i.e., the first) The first historical time window adjacent to each other (time window) (i.e., the first) (one time window), the first historical time window It is a time window The previous time window; determine the time window Second historical time windows K time windows apart (i.e., the first) (a time window), and determine its relationship with the first historical time window. The third historical time window that is K time windows apart (i.e., the first) (a total of 10 time windows), of which the second historical time window is... Located in the time window Previously, the third historical time window was located before the first historical time window. Then, the first similarity between the technological data of the second historical time window and the historical technological data of the third historical time window was determined. Specifically, for each time window... The scientific and technological data is used for feature encoding (the specific implementation can be found in the second vocabulary feature generation process described above) to obtain the current feature vector. ; Regarding the second historical time window Historical scientific and technological data are used for feature encoding (the specific implementation can be found in the second vocabulary feature generation process described above), resulting in historical feature vectors. , then calculate and The cosine similarity is denoted as the first similarity. Simultaneously, determine the second similarity between historical scientific and technological data from the first historical time window and historical scientific and technological data from the third historical time window. Specifically, for the first historical time window... Historical scientific and technological data are used for feature encoding (the specific implementation can be found in the second vocabulary feature generation process described above), resulting in historical feature vectors. ; Regarding the third historical time window Historical scientific and technological data are used for feature encoding (the specific implementation can be found in the second vocabulary feature generation process described above), resulting in historical feature vectors. , then calculate and The cosine similarity is denoted as the second similarity. Finally, the first similarity score is... Subtract the second similarity The degree of change is obtained. In practical applications, K can be determined according to different technical fields. K is related to the evolution cycle of the technical field. The longer the evolution cycle, the larger the value of K. If the scientific and technological data of the technical field shows periodicity, the value of K can also be the period length of the scientific and technological data cycle.
[0067] Applying the above embodiments, 1) Quantifying dynamic changes: By using cosine similarity, the semantic differences in scientific and technological data are transformed into quantitative indicators, enabling a quantifiable assessment of the "degree of technological change" and providing an objective basis for subsequent mutation detection and trend prediction. 2) Avoiding single historical bias: Three types of historical time windows are introduced (adjacent to the current time window, distant from the current time window, etc.). One, distance from the first historical time window (items), comparing "current and" "Step forward" and "Take a step forward and" The similarity difference between the "step before the previous step" avoids evaluation bias caused by relying solely on a single historical window, improving the stability of the change degree calculation. 3) Adapting to the domain rhythm: By adjusting The value can flexibly adapt to the pace of technological change in different fields, making the calculation of the degree of change more in line with the actual dynamics. 4) Accurately capture semantic differences: Based on feature vector cosine similarity calculation, it can accurately capture the deep semantic relationships of scientific and technological data, avoiding the shortcomings of traditional word frequency methods that cannot reflect semantic differences, and improving the accuracy of the degree of change calculation.
[0068] In some embodiments, see Figure 5The step "determining the first similarity between the scientific and technological data of the time window and the historical scientific and technological data of the second historical time window" can be achieved by executing the following steps 301-304: Step 301, determine the first importance of the hot technology fields of the time window and determine the second importance of the historical hot technology fields of the second historical time window; Step 302, extract the third vocabulary feature of each technical term in the time window, and weight the third vocabulary feature with the first importance as the weight value to obtain the fourth vocabulary feature; Step 303, extract the fifth vocabulary feature of each historical technical term in the second historical time window, and weight the fifth vocabulary feature with the second importance as the weight value to obtain the sixth vocabulary feature; Step 304, determine the feature similarity between the fourth vocabulary feature of the time window and the sixth vocabulary feature of the second historical time window, and use the feature similarity as the first similarity.
[0069] For step 301, as described in the above embodiments, the hot technical field is the technical field where the target technical term corresponding to the cluster center of the cluster is located. Therefore, for the hot technical fields within the time window, the word vectors of the target technical terms corresponding to the hot technical fields are extracted, and the norm of the word vector is calculated. If there is one hot technical field, the norm of the word vector corresponding to that hot technical field is taken as the first importance; if there are multiple hot technical fields, the average of the norms of the word vectors corresponding to the multiple hot technical fields is taken as the first importance. Similarly, the calculation process of the second importance can refer to the calculation process of the first importance, and will not be elaborated here.
[0070] For step 302, for each technical term within the time window, the third lexical feature (i.e., semantic vector, with dimension 1) of the technical term is extracted using the text vector model of the large language model. The third lexical feature is weighted according to the first importance weight to obtain the fourth lexical feature of the technical term.
[0071] For step 303, for each historical technical term within the second historical time window, the fifth lexical feature (i.e., semantic vector, with dimension 1) of the historical technical term is extracted using the text vector model of the large language model. The fifth lexical feature is weighted using the second importance weight to obtain the sixth lexical feature of the historical technical term.
[0072] For step 304, the average of all fourth-word features within the time window is calculated to obtain the overall first feature vector of the time window; the average of all sixth-word features within the second historical time window is calculated to obtain the overall second feature vector of the second historical time window; the cosine similarity between the first feature vector and the second feature vector is calculated, and the cosine similarity between the first feature vector and the second feature vector is used as the first similarity.
[0073] It should be noted that the calculation process for the second similarity can refer to the calculation process for the first similarity, and will not be repeated here.
[0074] Applying the above embodiments: 1) Focusing on core technology semantics: By weighting lexical features based on the importance of hot technology areas, the overall features of the window are made more closely aligned with the core technology content, avoiding interference from non-hot technology words and improving the targeting of similarity calculation. 2) Improving semantic matching accuracy: Calculating overall similarity based on weighted lexical features can more accurately reflect the semantic association between the two windows in the "core technology area," avoiding semantic bias caused by traditional methods relying on bare features. 3) Verifiable quantitative results: All calculations are based on quantifiable vectors and similarity indicators, providing objective basis for subsequent mutation event detection and technology trend prediction, avoiding subjective judgment errors.
[0075] Step 104: Based on the degree of change in each time window, determine the technological mutation events included in the scientific and technological data of multiple time windows.
[0076] For step 104, after calculating the degree of change for each time window, the technological change events included in the scientific and technological data of multiple time windows are determined based on the relationship between the degree of change for each time window and the degree of change threshold (which may be preset). Specifically, at least one of the following information for the technological change event is determined: the window number of the target time window involved in the technological change event, the hot technology fields before the technological change event, the hot technology fields after the technological change event, the degree of change of the target time window; etc.
[0077] In some embodiments, step 104, "determining the technological mutation events included in the scientific and technological data of multiple time windows based on the degree of change of each time window," can be achieved by performing the following steps: from multiple time windows, determining M target time windows whose degree of change is greater than a degree of change threshold and are temporally consecutive, where M is greater than or equal to a set value; and marking technological mutation events for the scientific and technological data of the M target time windows.
[0078] Here, a threshold for the degree of change (used to determine whether the degree of change in a single time window is significant) and M (used to determine whether the number of consecutive time windows meets the conditions for marking a technological mutation event) are preset. M is greater than or equal to a set value, which is an integer greater than 0, such as 2 or 3. M is greater than or equal to a set value, such as 3. The degree of change of multiple time windows is traversed, and a set of time windows that meet the following conditions is selected: (1) the degree of change of a single time window is greater than the threshold for the degree of change; (2) the time windows are consecutive in the time series, and the number of consecutive windows is greater than or equal to M. Thus, the M consecutive target time windows that meet the above conditions are recorded as a set of time windows. Finally, for each target time window, the scientific and technological data (including technical terms, hot technical fields, etc.) corresponding to the target time window are extracted, and technological mutation events are marked for the scientific and technological data corresponding to the target time window. For example, the marked content includes, but is not limited to: the time interval of the technological mutation event (the start number and end number of the target time window), the hot technical fields corresponding to the technological mutation event (including the hot technical fields before the technological mutation event and the hot technical fields after the technological mutation event), etc.
[0079] Applying the above embodiments: 1) Quantitative screening avoids subjective misjudgment: By setting a threshold for the degree of change, technological mutations are transformed from qualitative judgment to quantitative screening, avoiding the subjectivity of manual labeling. 2) Capturing continuous mutations and excluding isolated fluctuations: The requirement that the time window be continuous and the number greater than or equal to M ensures that technological mutation events are "continuous, trend-based changes" rather than isolated random fluctuations, improving the reliability of mutation events. 3) Providing precise analysis units: The labeled mutation events include time intervals and corresponding hot technology areas, providing a clear data foundation for subsequent technology trend prediction and enhancing the practicality of technology trend prediction.
[0080] Next, combine Figure 9 ( Figure 9 This is a flowchart illustrating the process of marking technology mutation events according to an embodiment of this application, detailing the process including: 1. Extracting feature vectors for each time window based on the importance of hot technology areas; 2. Calculating all window pairs (i.e., the current time window). Historical time windows before the current time window 1. Calculate the cosine similarity between the feature vectors of the target time window; 2. Calculate the difference rate of change (i.e., the degree of change mentioned above); 3. Determine the target time window for three consecutive difference rates of change greater than 0.3; 4. Mark the technical mutation event for the target time window.
[0081] Among them, the rate of change of difference It is calculated using the following formula (3):
[0082] in, For the current time window (i.e., the first) Feature vectors (within time windows); The second historical time window mentioned above (i.e., the first) Feature vectors (within time windows); The first historical time window mentioned above (i.e., the first) Feature vectors (within time windows); The aforementioned third historical time window (i.e., the first) The feature vectors of each time window. Differential rate of change. Used to measure the degree of deviation between the current time window and the historical time window (or historical baseline).
[0083] Detecting mutation events using technical methods, when the rate of change differs over three consecutive time windows. When the threshold is >0.3 (the threshold of 0.3 can be tested and verified during actual engineering implementation), it is determined to be a technical mutation signal.
[0084] Step 105: Based on technological upheaval events and the hot technology sectors in each time window, predict technological trends.
[0085] For step 105, after identifying the technological upheaval events and the hot technology areas for each time window, the technological trends are predicted by combining the technological upheaval events and the hot technology areas for each time window.
[0086] In some embodiments, see Figure 6 Step 105, "Predicting technology trends based on technology mutation events and hot technology areas in each time window," can be achieved by performing the following steps: Step 1051, extracting the hot technology time-series features of the hot technology areas in each time window, and extracting the mutation event time-series features of the technology mutation events; Step 1052, concatenating the hot technology time-series features and mutation event time-series features to obtain concatenated time-series features; Step 1053, extracting the hidden state features of the concatenated time-series features through a long short-term memory network, and performing attention processing on the hidden state features through an attention network to obtain attention features; Step 1054, fusing the hidden state features and attention features to obtain fused features, and predicting technology trends based on the fused features to obtain the technology trends.
[0087] For step 1051, firstly, for the hot technology fields of each time window, obtain relevant data for those hot technology fields (such as relevant technical terms in the hot technology fields, the importance of the hot technology fields). The process involves identifying the core keywords of hot technology fields, then extracting features from relevant data within each time window to obtain the temporal characteristics of hot technology fields for each time window. For each time window, mutation event data is acquired, including mutation markers (0 for no mutation, 1 for mutation), degree of change, etc. Features are then extracted from the mutation event data for each time window to obtain the temporal characteristics of mutation events for each time window.
[0088] For step 1052, the hot technology time series features and the mutation event time series features of each time window are concatenated to obtain the window time series features of each time window; the window time series features of multiple time windows are arranged in chronological order to obtain the time series feature sequence, which is the concatenated time series feature.
[0089] For step 1053, firstly, hidden state features of the concatenated temporal features are extracted using a Long Short-Term Memory (LSTM) network, and then attention features are applied to these hidden state features using an attention network to obtain attention features. Specifically, the concatenated temporal features are input into the LSTM network, which outputs hidden state features (a sequence of hidden state features, including window hidden state features for each time window). The attention network calculates the attention weight for each window hidden state feature; based on the attention weights of each window hidden state feature, the hidden state features of multiple windows are weighted and summed to obtain the attention features.
[0090] For step 1054, the hidden state features and attention features are fused to obtain the fused features. The fused features are then input into a fully connected layer for technology trend prediction to obtain the technology trend. Specifically, the fused features are mapped through the fully connected layer to obtain the technology trend. In practical applications, technology trends can be characterized by technology trend indicators. Therefore, technology trend prediction maps the fused features to quantifiable technology trend indicators. For example, technology trend indicators include, but are not limited to: a) Technology Importance I_pred: a continuous value, ranging from 0 to 1, representing the future importance of the technology field; b) Growth Rate G_pred: a continuous value (e.g., 0.15), representing the growth rate of the importance of the technology field (e.g., a 15% increase); c) Technology Field Size S_pred: a continuous value, representing the scale of scientific and technological data in the technology field, such as the number of documents / patents; d) Mutation Risk Probability P_pred: ranging from 0 to 1, representing the probability of a future technology mutation event; e) Confidence Interval [low, high]: representing the confidence range of the predicted value of the technology trend indicator, such as the confidence interval of I_pred [0.88, 0.95]).
[0091] Steps 1053 and 1054 can be implemented using a trained hybrid LSTM-Attention model.
[0092] Applying the above embodiments, 1) Multi-source information fusion: By splicing together the temporal features of hot technologies and the temporal features of abrupt events, the "conventional trends" and "abrupt disturbances" of technological development are integrated, avoiding the loss of information from a single feature. 2) Temporal dependency capture: LSTM extracts hidden state features, effectively capturing the long-term temporal dependencies of technological trends (such as the impact of preceding windows on subsequent ones), improving the temporal consistency of predictions. 3) Key event highlighting: The attention network weights the hidden state features, highlighting time points containing abrupt events or important technological changes, enhancing the relevance of the hidden state features. 4) Improved prediction accuracy: By fusing hidden state features and attention features, the impact of temporal dependencies and key events is integrated, ultimately improving the accuracy and reliability of technological trend predictions.
[0093] Next, combine Figure 10 ( Figure 10 This is a flowchart illustrating the process of predicting technology trends provided in this application embodiment. The process of predicting technology trends is described in detail, including: 1. Identifying hot technology fields; 2. Identifying technology mutation events; 3. Merging hot technology fields and technology mutation events to obtain a fused event signal (i.e., the aforementioned fused feature); 4. Training a hybrid LSTM-Attention model; 5. Using the trained hybrid LSTM-Attention model, predicting technology trends based on the fused event signal.
[0094] It should be noted that hot technology sectors typically appear frequently over a sustained period, and their data exhibits continuity and trends, making them suitable as fundamental input for technology trend prediction. Technological upheavals, on the other hand, occur suddenly and have a significant impact on the direction of technological development over discrete time periods. Their data is characterized by suddenness and discontinuity, and can provide corrective signals for trend prediction. By fusing hot technology sectors and technological upheavals, and using a trained hybrid LSTM-Attention model based on the fused event signals for technology trend prediction, prediction accuracy can be improved. Furthermore, the hybrid LSTM-Attention model is also trained based on the fused results of hot technology sectors and technological upheavals, further enhancing the prediction accuracy of the trained LSTM-Attention model.
[0095] By applying the embodiments described above, firstly, technological data from multiple time windows are acquired and technical terms are extracted. Clustering is then used to automatically identify hot technology areas within each time window, improving the efficiency of technological trend perception and preventing the omission of important information by humans. Secondly, by determining the degree of change in technological data for each time window compared to historical time windows, technological changes are quantified and technological mutation events are accurately identified, solving the problem of the difficulty for humans to accurately judge technological mutations. Finally, technological trends are predicted based on technological mutation events and hot technology areas, integrating key information on technological dynamics (i.e., technological mutation events and hot technology areas), thus improving the accuracy of technological trend prediction. In summary, the embodiments of this application achieve automated and precise perception of technological trends, improving the accuracy and efficiency of technological trend perception.
[0096] The following description continues to illustrate the exemplary structure of the scientific and technological data processing apparatus 555 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2 As shown, the software modules in the technology data processing device 555 stored in the memory 550 may include: an acquisition module 5551, used to acquire technology data for each of multiple time windows, the technology data including multiple technical terms; a clustering module 5552, used to cluster the multiple technical terms of each time window to obtain clusters, and to take the technical field represented by the technical term corresponding to the cluster center of the cluster as the hot technical field of the time window; a first determination module 5553, used to determine the degree of change of the technology data of each time window compared with the historical technology data of historical time windows; a second determination module 5554, used to determine the technological mutation events included in the technology data of the multiple time windows based on the degree of change of each time window; and a prediction module 5555, used to predict the technology trend based on the technological mutation events and the hot technical fields of each time window.
[0097] In some embodiments, the acquisition module 5551 is further configured to: acquire a first window size and a window adjustment coefficient of an initial time window before acquiring the scientific and technological data of each of the multiple time windows; acquire the time change of the i-th time window relative to the initial time window, and determine a second window size of the i-th time window based on the first window size, the window adjustment coefficient, and the time change, and determine an i-th time window having the second window size; traverse i to obtain the multiple time windows, wherein the multiple time windows are sequential in time, the starting time point of the first time window is the starting time point of the initial time window, i is an integer greater than 0 and less than or equal to N, and N is the number of time windows.
[0098] In some embodiments, the clustering module 5552 is further configured to: determine the importance of each technical term to a target technical description statement for each technical term in the time window, wherein the scientific and technological data in the time window includes the target technical description statement and the target technical description statement includes the technical term; extract a first lexical feature for each technical term, and weight the first lexical feature of the technical term with the importance as the weight value of the technical term to obtain a second lexical feature of the technical term; and cluster multiple second lexical features to obtain the cluster.
[0099] In some embodiments, the clustering module 5552 is further configured to: identify multiple technical entities from the scientific and technological data and encode the multiple technical entities to obtain technical entity features before clustering multiple technical terms in the time window to obtain clusters; extract multiple technical description statements including at least one of the technical entities from the scientific and technological data and encode each technical description statement to obtain statement features; determine the feature similarity between the statement features and the technical entity features of each technical description statement; and select the technical description statements whose feature similarity satisfies the similarity condition as the target technical description statements.
[0100] In some embodiments, the clustering module 5552 is further configured to encode each technical description statement using a text vector model of a large language model to obtain statement features; the clustering module 5552 is further configured to perform attention processing on the technical terms using an attention processing model of the large language model to obtain the importance of the technical terms to the target technical description statement; the clustering module 5552 is further configured to perform text vectorization processing on the technical terms using the large language model to obtain the first lexical features of the technical terms.
[0101] In some embodiments, the first determining module 5553 is further configured to determine a first historical time window adjacent to the time window, and to determine a second historical time window and a third historical time window; wherein the second historical time window is K time windows away from the time window, the third historical time window is located before the first historical time window and is K time windows away from the first historical time window, the historical time window includes the first historical time window, the second historical time window and the third historical time window, and K is an integer greater than 1; determine a first similarity between the scientific and technological data of the time window and the historical scientific and technological data of the second historical time window, and determine a second similarity between the historical scientific and technological data of the first historical time window and the historical scientific and technological data of the third historical time window; subtract the second similarity from the first similarity to obtain the degree of change.
[0102] In some embodiments, the first determining module 5553 is further configured to determine a first importance of the hot technology field in the time window, and determine a second importance of the historical hot technology field in the second historical time window; extract a third lexical feature of each technology term in the time window, and weight the third lexical feature with the first importance as a weight value to obtain a fourth lexical feature; extract a fifth lexical feature of each historical technology term in the second historical time window, and weight the fifth lexical feature with the second importance as a weight value to obtain a sixth lexical feature; determine the feature similarity between the fourth lexical feature of the time window and the sixth lexical feature of the second historical time window, and use the feature similarity as the first similarity.
[0103] In some embodiments, the second determining module 5554 is further configured to determine M target time windows from the plurality of time windows that have a change degree greater than a change degree threshold and are temporally consecutive, wherein M is greater than or equal to a set value; and to mark the technological mutation event for the technological data of the M target time windows.
[0104] In some embodiments, the prediction module 5555 is further configured to extract the hot technology time-series features of the hot technology fields in each time window, and extract the mutation event time-series features of the technology mutation events; concatenate each hot technology time-series feature and the mutation event time-series feature to obtain concatenated time-series features; extract the hidden state features of the concatenated time-series features through a long short-term memory network, and perform attention processing on the hidden state features through an attention network to obtain attention features; fuse the hidden state features and the attention features to obtain fused features, and perform technology trend prediction based on the fused features to obtain the technology trend.
[0105] It should be noted that the description of the apparatus embodiments in this application is similar to the description of the method embodiments described above, and has similar beneficial effects as the method embodiments, so it will not be repeated here. Any technical details not covered in the data processing apparatus provided in the embodiments of this application can be understood based on the description of the technical details in the above method embodiments.
[0106] This application also provides a computer program product, which includes computer-executable instructions or a computer program stored in a computer-readable storage medium. The processor of an electronic device reads the computer-executable instructions or computer program from the computer-readable storage medium and executes the computer-executable instructions or computer program, causing the electronic device to perform the scientific and technological data processing method provided in this application.
[0107] This application also provides a computer-readable storage medium storing computer-executable instructions or computer programs. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the scientific and technological data processing method provided in this application.
[0108] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0109] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.
[0110] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).
[0111] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located in one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0112] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A method for processing scientific and technological data, characterized in that, The method includes: Acquire scientific and technological data for each of multiple time windows, the scientific and technological data including multiple technical terms; For each time window, multiple technical terms within the time window are clustered to obtain clusters, and the technical fields represented by the technical terms corresponding to the cluster centers of the clusters are taken as the hot technical fields of the time window. For each time window, determine the degree of change in the scientific and technological data of that time window compared to the historical scientific and technological data of historical time windows; Based on the degree of change in each of the time windows, the technological mutation events included in the scientific and technological data of the multiple time windows are determined; Based on the aforementioned technological upheaval events and the hot technological fields in each of the aforementioned time windows, technological trends are predicted.
2. The method as described in claim 1, characterized in that, Before acquiring the scientific and technological data for each of the multiple time windows, the method further includes: Obtain the first window size and window adjustment factor of the initial time window; Obtain the time change of the i-th time window relative to the initial time window from the plurality of time windows, and determine the second window size of the i-th time window based on the first window size, the window adjustment coefficient and the time change, and determine the i-th time window with the second window size; Traverse i to obtain the plurality of time windows, which are consecutive in time. The starting time of the first time window is the starting time of the initial time window. i is an integer greater than 0 and less than or equal to N, and N is the number of time windows.
3. The method as described in claim 1, characterized in that, The clustering of multiple technical terms within the time window to obtain clusters includes: For each technical term in the time window, determine the importance of the technical term to the target technical description statement, wherein the scientific and technological data in the time window includes the target technical description statement, and the target technical description statement includes the technical term; For each of the technical terms, a first lexical feature of the technical term is extracted, and the first lexical feature of the technical term is weighted with the importance as the weight value of the technical term to obtain a second lexical feature of the technical term; Clustering is performed on multiple second vocabulary features to obtain the clusters.
4. The method as described in claim 3, characterized in that, Before clustering multiple technical terms within the time window to obtain clusters, the method further includes: Multiple technical entities are identified from the scientific and technological data, and the multiple technical entities are encoded to obtain technical entity features; From the scientific and technological data, extract multiple technical description statements that include at least one of the technical entities, and encode each technical description statement to obtain statement features; For each technical description statement, determine the feature similarity between the statement features and the technical entity features of the technical description statement; The technical description statement whose feature similarity satisfies the similarity condition among the plurality of technical description statements is taken as the target technical description statement.
5. The method as described in claim 4, characterized in that, The process of encoding each of the technical description statements to obtain statement features includes: Each of the technical description statements is encoded using the text vector model of a large language model to obtain statement features; Determining the importance of the technical terms to the target technical description statement includes: The technical terms are subjected to attention processing using the attention processing model of the large language model to obtain the importance of the technical terms to the target technical description statement; The extraction of the first lexical features of the technical terms includes: The technical terms are vectorized using the large language model to obtain the first lexical features of the technical terms.
6. The method as described in claim 1, characterized in that, Determining the degree of change of the scientific and technological data in the time window compared to the historical scientific and technological data in the historical time window includes: Determine a first historical time window adjacent to the stated time window, and determine a second historical time window and a third historical time window; Wherein, the second historical time window is K time windows away from the first historical time window, the third historical time window is located before the first historical time window and is K time windows away from the first historical time window, the historical time window includes the first historical time window, the second historical time window and the third historical time window, and K is an integer greater than 1; Determine the first similarity between the scientific and technological data of the time window and the historical scientific and technological data of the second historical time window, and determine the second similarity between the historical scientific and technological data of the first historical time window and the historical scientific and technological data of the third historical time window; The degree of change is obtained by subtracting the second similarity from the first similarity.
7. The method as described in claim 6, characterized in that, The determination of the first similarity between the scientific and technological data of the time window and the historical scientific and technological data of the second historical time window includes: Determine the first importance of the hot technology field in the time window, and determine the second importance of the historical hot technology field in the second historical time window; Extract the third lexical feature of each of the technical terms in the time window, and weight the third lexical feature with the first importance as the weight value to obtain the fourth lexical feature; Extract the fifth lexical feature of each historical technical term in the second historical time window, and weight the fifth lexical feature with the second importance as the weight value to obtain the sixth lexical feature; Determine the feature similarity between the fourth lexical feature of the time window and the sixth lexical feature of the second historical time window, and use the feature similarity as the first similarity.
8. The method as described in claim 1, characterized in that, The determination of technological abrupt events in the scientific and technological data of the multiple time windows based on the degree of change in each time window includes: From the plurality of time windows, determine M target time windows whose degree of change is greater than a change degree threshold and are consecutive in time, wherein M is greater than or equal to a set value; For the technological data of the M target time windows, the technological mutation events are marked.
9. The method as described in claim 1, characterized in that, The predicted technology trends, based on the technological mutation events and the hot technology fields of each time window, include: Extract the hot technology time series features of the hot technology field in each time window, and extract the mutation event time series features of the technology mutation events; The temporal features of each of the hotspot technologies and the temporal features of the mutation events are concatenated to obtain the concatenated temporal features. The hidden state features of the spliced temporal features are extracted by a long short-term memory network, and the hidden state features are then processed by an attention network to obtain attention features. The hidden state features and the attention features are fused to obtain fused features, and the technology trend is predicted based on the fused features.
10. A device for processing scientific and technological data, characterized in that, The device includes: The acquisition module is used to acquire scientific and technological data for each of the multiple time windows, the scientific and technological data including multiple technical terms; The clustering module is used to cluster multiple technical terms in each time window to obtain clusters, and to take the technical fields represented by the technical terms corresponding to the cluster centers of the clusters as the hot technical fields of the time window. The first determining module is used to determine, for each time window, the degree of change of the scientific and technological data of the time window compared with the historical scientific and technological data of the historical time window; The second determining module is used to determine the technological mutation events included in the scientific and technological data of the plurality of time windows based on the degree of change of each time window; The prediction module is used to predict technology trends based on the technological mutation events and the hot technology areas in each time window.
Citation Information
Patent Citations
Method and system for analyzing and predicting theme trend of scientific and technical literature
CN120068882A
Prospective power technology prediction method based on topic extraction and time modeling
CN120804313A
Emergency sensing method including semantic extraction
CN120911604A