Multi-source heterogeneous astronomical data management method and device and medium
By constructing an astronomical knowledge graph for semantic alignment, the problem of inconsistent data formats of different astronomical equipment is solved, unified management and rapid query of multi-source heterogeneous astronomical data is realized, and fusion analysis of astronomical data is supported.
Patent Information
- Application Number
- CN202510765257.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When different astronomical observation equipment uses the FITS data format, there are problems such as different synonyms and different names and different meanings of the same name, which makes it difficult to achieve unified management and efficient query of massive astronomical data, limiting the development of astronomical data fusion analysis.
By obtaining archive task instructions for new observation data and pre-constructed astronomical knowledge graphs, the metadata of file headers is parsed, and based on semantic alignment technology, the metadata is uniformly managed into a designated database, and the astronomical knowledge graph is used for semantic alignment, so as to achieve unified representation and management of multi-source heterogeneous astronomical data.
It realizes unified management of astronomical data, avoids query failure caused by different synonyms or different meanings of the same name, supports rapid query and fusion analysis, and promotes the development of astronomy.
Smart Images

Figure CN120295970A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data management, and particularly to a method, device, and medium for managing multi-source heterogeneous astronomical data. Background Art
[0002] In the rapid development of astronomy technology, it is crucial for promoting the further development of astronomy to fuse and analyze different types of astronomical data from all over the world to form a larger virtual telescope. Currently, in order to achieve the fusion analysis of massive astronomical data, the FITS (Flexible Image Transport System) data format is used as the unified standard for data transmission and exchange between different observatories around the world.
[0003] However, although different astronomical observation devices use the same FITS data format, different astronomical observation devices can customize the file header of the FITS file, and the data in the FITS file is closely related to the file header. In the file header, there may be problems of semantic heterogeneity such as synonymous but different names, or the same name but different meanings for different astronomical observation devices. For example, for the right ascension, some astronomical observation devices describe it as "MC_RA", while some describe it as "RA". In addition, during the long-term on-orbit operation of astronomical observation devices, the same thing may generate new and old names, making it difficult to achieve a unified representation of astronomical data, thus preventing the efficient management and query of massive astronomical data and restricting the development of the fusion analysis of massive astronomical data.
[0004] Therefore, how to uniformly manage astronomical data and achieve rapid query of multi-source heterogeneous astronomical data is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, one aspect of this application provides a method for managing multi-source heterogeneous astronomical data, and the method includes: Obtain the archiving task instruction of new observation data and a pre-constructed astronomy knowledge graph; Parse the archiving task instruction; and establish a communication connection with the data source device that collects the new observation data according to the parsing result; Scan the file header of the new observation data; wherein, the file header includes metadata for describing information related to the new observation data; Based on the astronomy knowledge graph, perform a semantic alignment operation on the metadata in the file header; and store the semantic alignment result in a specified database.
[0006] Optionally, performing a semantic alignment operation on the metadata in the file header based on the astronomy knowledge graph includes: Extract the first file header scanned from each of the data source devices; Determine whether there is a historical alignment result in the specified database for semantic alignment of the metadata in the first file header; If there is, replace the metadata in the file header according to the historical alignment result; If not, based on the astronomy knowledge graph, perform semantic alignment on the metadata in the first file header through a specified model; and replace the other metadata in the file header except the first file header according to the current alignment result.
[0007] Optionally, before replacing the other metadata in the file header except the first file header according to the current alignment result, it further includes: If an alignment modification instruction from the user is obtained within a preset time period; Perform secondary semantic alignment on the metadata in the first file header according to the alignment modification instruction; and use the secondary semantic alignment result as the current alignment result; Collect the current alignment result to obtain a training data set; Train the specified model through the training data set.
[0008] Optionally, the method for managing multi-source heterogeneous astronomical data further includes: Obtain a query statement of the user for multi-source heterogeneous astronomical data; Parse the query statement to determine the semantic information and included word units of the query statement; Construct a syntax tree for describing the hierarchical relationship between the word units according to the semantic information; wherein, the syntax tree includes leaf nodes and a root node, and the leaf nodes and the root node are the word units; Replace the leaf nodes with target fields in the astronomy knowledge graph to obtain a target query statement; Obtain a query result of multi-source heterogeneous astronomical data in the specified database through the target query statement.
[0009] Optionally, the parsing the query statement to determine the semantic information and included word units of the query statement includes: Based on predefined lexical rules, segment the query statement to obtain the word units; Call a syntax analyzer to perform syntax analysis on the word units, and perform semantic analysis on the word units through a semantic analyzer to obtain the semantic information.
[0010] Optionally, constructing an astronomy knowledge graph includes: Obtain astronomical data sets; Extract entity fields and relationship fields from the astronomical data sets; Construct the entity fields and the relationship fields into triples; Generate an astronomical knowledge graph based on the triples.
[0011] Optionally, parsing the archiving task instruction includes: Extract the task files in the archiving task instruction; Parse the task files to determine the access parameters of the data source device and the access logic of the file headers.
[0012] Another aspect of the present application provides a management device for multi-source heterogeneous astronomical data, and the device includes: An acquisition module, configured to acquire an archiving task instruction of new observation data and a pre-constructed astronomical knowledge graph; A communication connection module, configured to parse the archiving task instruction; and establish a communication connection with a data source device that acquires the new observation data according to the parsing result; A scanning module, configured to scan the file headers of the new observation data; wherein, the file headers include metadata for describing information related to the new observation data; A semantic alignment module, configured to perform a semantic alignment operation on the metadata in the file headers based on the astronomical knowledge graph; and store the semantic alignment result in a specified database.
[0013] Another aspect of the present application provides a management device for multi-source heterogeneous astronomical data, including a memory and a processor, where a computer program that can run on the processor is stored on the memory, and when the processor executes the program, the steps of the management method for multi-source heterogeneous astronomical data are implemented.
[0014] Another aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the management method for multi-source heterogeneous astronomical data are implemented.
[0015] The beneficial effects produced by the management method, device and medium for multi-source heterogeneous astronomical data provided by the present application are as follows: Based on the astronomical knowledge graph, semantic alignment is performed on the scanned file header metadata, realizing unified management of astronomical data, providing support for subsequent rapid query of multi-source heterogeneous astronomical data, avoiding query failures caused by semantic heterogeneity such as different names for the same meaning or the same name for different meanings, and being able to obtain query results with different names for the same meaning or the same name for different meanings simultaneously, promoting the development of astronomical fusion analysis. Description of the Drawings
[0016] Figure 1Schematic flowchart of a method for managing multi-source heterogeneous astronomical data provided by an embodiment of the present application; Figure 2 Schematic structural diagram of a system for managing multi-source heterogeneous astronomical data provided by an embodiment of the present application; Figure 3 Schematic flowchart of a method for managing multi-source heterogeneous astronomical data provided by another embodiment of the present application; Figure 4 Schematic structural diagram of a device for managing multi-source heterogeneous astronomical data provided by an embodiment of the present application; Figure 5 Schematic structural diagram of a device for managing multi-source heterogeneous astronomical data provided by another embodiment of the present application.
[0017] Reference numerals are as follows: 1 is a management system, 2 is a data source device, 3 is an interaction unit, 10 is an astronomy knowledge graph management unit, 11 is a data source device connection unit, 12 is a semantic alignment unit, 13 is a specified database, 14 is a query unit, 40 is an acquisition module, 41 is a communication connection module, 42 is a scanning module, 43 is a semantic alignment module, 50 is a memory, 51 is a processor, 52 is a display screen, 53 is an input / output interface, 54 is a communication interface, 55 is a power supply, 56 is a communication bus, 501 is a computer program, 502 is an operating system, 503 is data. Detailed implementation manners
[0018] The terms used in the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "said", and "the" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0019] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".
[0020] Figure 1 Schematic flowchart of a method for managing multi-source heterogeneous astronomical data provided by an embodiment of the present application, as Figure 1 shown, the method includes: S10: Obtain the archiving task instruction for the new observation data and the pre-constructed astronomy knowledge graph; S11: Parse the archiving task instruction; and establish a communication connection with the data source device for collecting the new observation data according to the parsing result; Figure 2 The following is a schematic structural diagram of a management system for multi-source heterogeneous astronomical data provided by an embodiment of the present application. In an alternative embodiment, the management method for multi-source heterogeneous astronomical data provided by the present application can be applied to this management system, that is, it can be implemented through this management system. As Figure 2 shown, this management system 1 includes: an astronomy knowledge graph management unit 10, a data source device connection unit 11, and a semantic alignment unit 12.
[0021] In a specific embodiment, the astronomy knowledge graph management unit 10 is used to uniformly manage information such as terms and concepts in the astronomy field in the form of a knowledge graph. This unit supports operations such as adding, deleting, modifying, and querying information such as terms and concepts in the astronomy field, so as to maintain the subordinate relationships between different concepts and terms.
[0022] In an alternative embodiment, the user can initiate an archiving task instruction for new observation data to the management system 1 through the interaction unit 3. The data source device connection unit 11 in the management system 1 obtains this archiving task instruction and parses the archiving task instruction, so as to establish a communication connection with the data source device 2 according to the parsing result.
[0023] It can be understood that in order to establish a communication connection with the data source, when parsing the archiving task instruction, it is necessary to obtain the access parameters of the data source device 2, for example, parameter information such as the IP address, communication port, username, and password of the data source device 2.
[0024] It should be noted that the newly generated observation data can be archived in real time, or the newly generated observation data can be archived once every preset period. This application does not make any limitations in this regard. In addition, it should also be noted that the data source device 2 is an observatory observation device, and this application does not make any limitations on the type, structure, model, etc. of the observatory observation device. In a specific embodiment, there are multiple data source devices 2, and they can be data source devices 2 in different regions. This application does not make any limitations on the number and specific regions of the data source devices 2.
[0025] S12: Scan the file header of the new observation data; wherein, the file header includes metadata for describing information related to the new observation data; Further, after establishing a communication connection with the data source device 2, the file header of the new observation data can be scanned to obtain a data file. It should be noted that different data files are composed of the new observation data itself and the file header used to describe the relevant information of the new observation data. It can be understood that astronomical observation data is generated and observed in real time, and the data volume is very large. If all astronomical observation data is uniformly managed, it will cause huge resource overhead. In addition, the massive astronomical observation data will also cause huge transmission pressure during the transmission process, and will seriously affect the efficiency of unified management. Therefore, in an alternative embodiment, only the file headers of astronomical observation data are scanned and uniformly managed.
[0026] Different data source devices 2 can customize the format of the file header. The metadata in the file header (i.e., the data used to describe the new observation data itself) can include, but is not limited to, the observation time, the observation device number, the observation module prefix, and the exposure parameters, etc.
[0027] For example, different data source devices 2 can customize different FITS file header formats, but the data captured and generated by the same data source device 2 has the same file header format, but the corresponding new observation data itself and the metadata in the file header (such as exposure parameters, etc.) are different.
[0028] S13: Based on the astronomy knowledge graph, perform semantic alignment operations on the metadata in the file header; and store the semantic alignment results in a specified database.
[0029] In the unified management of astronomical data, in order to avoid semantic heterogeneity such as different names for the same meaning or the same name for different meanings, Figure 2 As shown in the semantic alignment unit 12, based on the astronomy knowledge graph pre-constructed in the astronomy knowledge graph management unit 10, perform semantic alignment operations on the metadata in the scanned file header. For example, for "right ascension", some data source devices 2 describe it with MC_RA, some with RA, while in the pre-constructed astronomy knowledge graph, it is described through the "right ascension" field. Therefore, after semantic alignment of the metadata fields through step S13, they can all be described as "right ascension".
[0030] Further, store the semantic alignment results in the specified database 13. In an alternative embodiment, the semantic alignment results can be stored in the form of matching key-value pairs, where the matching key-value pairs are composed of the metadata field name and the knowledge graph node name. In another alternative embodiment, the specified database 13 is a vector database. Correspondingly, the semantic alignment results (i.e., the matching key-value pairs) are persistently stored in the vector database one by one for subsequent quick query.
[0031] For example, for each semantic alignment result, a piece of information is generated and sent to the Kafka-based message caching device. In the generated message, the corresponding fields in the FITS file header will be replaced with the results after semantic alignment. For example, the description of the "right ascension" field in the original FITS file header will be replaced with the concept "TERM_right_ascension". Further, through the message caching device, the received messages are processed one by one and persisted into the Milvus-based vector database. By introducing technical means such as message caching devices and vector databases, the control of message traffic is realized, the peak is shaved and the valley is filled, the stability of the service is ensured, and the efficiency and performance of the system are improved.
[0032] Thus, the method for managing multi-source heterogeneous astronomical data provided by the embodiments of the present application performs semantic alignment on the scanned file header metadata based on the astronomical knowledge graph, realizes the unified management of astronomical data, provides support for the subsequent rapid query of multi-source heterogeneous astronomical data, avoids query failures caused by semantic heterogeneity such as different names for the same meaning or the same name for different meanings, and can simultaneously query and obtain query results with different names for the same meaning or the same name for different meanings, promoting the development of astronomical fusion analysis.
[0033] In an alternative embodiment, based on the astronomical knowledge graph, performing a semantic alignment operation on the metadata in the file header includes: Extracting the first file header scanned from each data source device; Determining whether there is a historical alignment result for performing a semantic alignment operation on the metadata in the first file header in the specified database; If it exists, replacing the metadata in the file header according to the historical alignment result; If it does not exist, based on the astronomical knowledge graph, performing a semantic alignment operation on the metadata in the first file header through the specified model; and replacing the metadata in the other file headers except the first file header according to the current alignment result.
[0034] In a specific embodiment, it can be understood that the astronomical data generated by the same astronomical observation device has the same file header format. Therefore, in order to save resources, only the first file read needs to be processed. Specifically, the first file headers scanned from different data source devices 2 are extracted.
[0035] Furthermore, Figure 2 The semantic alignment unit 12 shown determines whether the same semantic alignment operation has been processed in the historical archiving task. Specifically, a query is performed in the specified database 13. When there is a historical alignment result, the historical alignment result can be directly applied to all the file headers scanned in the current archiving task. Specifically, according to the historical alignment result, all the metadata in the file header is replaced.
[0036] Of course, if there is no historical alignment result in the specified database 13, the specified model can be called to perform semantic alignment on the metadata in the first extracted file header. The specified model can be a deep learning-based model or a large language model based on the Transformer architecture, and this application does not make any limitations in this regard.
[0037] After performing semantic alignment on the first file header, all the current alignment results are applied to the metadata of the remaining other file headers, that is, according to the current alignment results, the metadata of the file headers except the first file header are replaced. For example, "MC_RA" representing right ascension in the file header is replaced with "TERM_right_ascension". Similarly, in a specific embodiment, the current alignment results are stored in the specified database 13.
[0038] In an alternative embodiment, Figure 2 the semantic alignment unit 12 shown performs semantic alignment on the first file header scanned from the heterogeneous data source adaptation device using a deep learning model. It should be noted that according to the characteristics of astronomical data, different telescopes can customize the file header, but the data generated by these telescopes all use the same data format.
[0039] Furthermore, the aligned results are applied to the process of traversing all the remaining subsequent files, and in all concept alignment operations, a corresponding message will be generated for the metadata of each file. The semantic alignment operation refers to re-labeling the scanned metadata using the information stored in the astronomical knowledge graph management unit. For example, in some FITS file headers, the column name prefix_tel is used to represent a certain telescope, and in FITS file headers from other sources, telescope is used to represent the same meaning, but in the archiving of the embodiments of this application, TERM_telescope can be used for unified representation.
[0040] Thus, through semantic alignment operations, the unified representation and management of multi-source heterogeneous data are achieved, facilitating subsequent rapid query of multi-source heterogeneous data.
[0041] Based on the above embodiments, as an alternative embodiment, before replacing the metadata of the file headers except the first file header according to the current alignment results, it further includes: If an alignment modification instruction from the user is obtained within a preset time period; Performing secondary semantic alignment on the metadata of the first file header according to the alignment modification instruction; and using the secondary semantic alignment result as the current alignment result; Collecting the current alignment results to obtain a training dataset; Train a specified model using a training dataset.
[0042] In a specific embodiment, to further improve the reliability of unified management of astronomical data, the user can view the semantic alignment result of each time in real time through the interaction unit 3. Therefore, after performing a semantic alignment operation on the first file header using the specified model, wait for a preset duration. If within this preset duration, the user issues an alignment modification instruction, that is, the user is not satisfied with the current semantic alignment result, thus, the semantic alignment accuracy and reliability can be improved through manual intervention.
[0043] In a specific embodiment, obtain the alignment modification instruction output by the user within the preset duration, and perform a secondary semantic alignment on the metadata of the first file header according to the alignment modification instruction. Further, apply the secondary semantic alignment result to the metadata in the remaining file headers.
[0044] For example, when performing the first semantic alignment operation using the specified model, the metadata A of the first file header is replaced with field B. However, after the user views the first alignment result and is not satisfied, and wants to replace field B with field C. At this time, input an alignment modification instruction including field B and field C through the Figure 2 shown interaction unit 3 to perform a secondary semantic alignment on the metadata of the first file header.
[0045] In an alternative embodiment, to ensure that the subsequent semantic alignment can meet the user's expectations, collect each obtained current alignment result as a training dataset. And train the specified model using this training dataset. Thus, in the next semantic alignment task, the specified model can more accurately align the results that satisfy the user.
[0046] Thus, by introducing technical means such as a recommendation model and an expert system, the optimization and improvement of semantic conversion and pattern mapping are realized, supporting more types of data sources and scenarios, and being able to accept new concepts and information. Continuously train the specified model using the constructed training dataset to update and enhance the capabilities of the specified model.
[0047] Figure 3 It is a flowchart of a method for managing multi-source heterogeneous astronomical data provided by another embodiment of the present application. In an alternative embodiment, as Figure 3 shown, the method for managing multi-source heterogeneous astronomical data provided by the present application further includes: S30: Obtain a query statement of the user for multi-source heterogeneous astronomical data; S31: Parse the query statement to determine the semantic information and included word units of the query statement; It can be understood that through the above embodiments, the specified database 13 includes a large number of semantic alignment results, that is, a large number of matching key-value pairs. Based on this specified database 13, queries for multi-source heterogeneous astronomical data can be performed.
[0048] Specifically, through the interaction unit 3 as Figure 2 shown, the query statement output by the user is obtained. In an optional embodiment, the query statement can be a DSL (Domain Specific Language) query statement. Further, the DSL query statement is parsed to determine the word units (tokens) included in the DSL query statement and the semantic information of the DSL query statement. It should be noted that the word unit token can be but is not limited to keywords, identifiers, constants, operators, etc., and this application does not make any limitations in this regard.
[0049] S32: According to the semantic information, construct a syntax tree for describing the hierarchical relationship between word units; wherein, the syntax tree includes leaf nodes and a root node, and the leaf nodes and the root node are word units; S33: Replace the leaf nodes with target fields in the astronomical knowledge graph to obtain a target query statement; S34: Through the target query statement, obtain the query result of multi-source heterogeneous astronomical data in the specified database.
[0050] Further, Figure 2 the query unit 14 as shown gradually constructs a syntax tree according to the semantic information. The syntax tree is a tree-shaped data structure, and its nodes represent the syntax structures in the code, and the parent-child relationship between the nodes reflects the hierarchical relationship between the syntax structures. That is, the syntax tree can be used to describe the hierarchical relationship between word units, and the word units serve as the leaf nodes and the root node of the syntax tree. For example, for the DSL query statement "(3 + 5) * 2", the root node of the syntax tree can be the multiplication operator "*", its left leaf node can be "3 + 5", and its right leaf node can be the constant "2".
[0051] It should be noted that at least one of all the leaf nodes of the syntax tree exists in the already constructed astronomical knowledge graph. According to the hierarchical relationship between the nodes in the astronomical knowledge graph, the leaf nodes constructed from the original DSL query statement can be replaced with astronomical terms, thereby completing the reconstruction of the DSL query statement.
[0052] Specifically, based on the syntax tree, the leaf nodes in the DSL query statement are replaced with target fields in the astronomical knowledge graph to obtain a target query statement. Thus, through the target query statement, the query result of multi-source heterogeneous astronomical data can be obtained in the specified database 13. For the convenience of understanding, an example will be given below.
[0053] For example, for "telescope", in the specified database 13, there are matching key-value pairs replacing metadata A1, metadata A2, and metadata A3 with "telescope". Metadata A1, metadata A2, and metadata A3 respectively come from different astronomical observation devices, and the target query statement reconstructed through the syntax tree points to the query for "telescope". Thus, all "telescopes" corresponding to different astronomical observation devices can be queried.
[0054] In an alternative embodiment, the query statement is parsed to determine the semantic information and included word units of the query statement, including: Based on predefined lexical rules, the query statement is segmented to obtain word units; The lexical analyzer is called to perform syntax analysis on the word units, and the semantic analyzer is used to perform semantic analysis on the word units to obtain semantic information.
[0055] In a specific embodiment, the user can submit a query statement dedicated to the astronomy field through the interaction unit 3 as shown in Figure 2 . In an alternative embodiment, the query statement can be a DSL statement. After the user submits the DSL query statement, it is processed through four stages: lexical analysis, syntax analysis, semantic analysis, and construction of the syntax tree of the DSL query statement in sequence to construct a new DSL query statement.
[0056] First, lexical analysis refers to segmenting the query statement based on predefined lexical rules to obtain word units. Specifically, the lexical analyzer scans and segments the DSL query statement in the order of the character stream so as to segment the original DSL query statement into individual word units (Tokens).
[0057] Among them, the predefined lexical rules refer to Token segmentation rules. In a specific embodiment, Tokens can include but are not limited to keywords, identifiers, constants, operators, and delimiters. For example, for the code "int num = 10", the lexical analyzer will segment it into multiple word units such as "int" (keyword), "num" (identifier), "=" (operator), "10" (constant), and ";" (delimiter). It should be noted that in an alternative embodiment, the lexical analyzer can be implemented using a Finite State Automaton (FSA).
[0058] Further, syntax analysis is performed, that is, a syntax analyzer is called to perform syntax analysis on the lexical units. In a specific embodiment, the Tokens output by the lexical analyzer are used as the input end of the syntax analyzer, and the Tokens are analyzed according to the syntax rules of the DSL to check whether the code conforms to the syntax rules. In an alternative embodiment, the syntax rules of the DSL can be described using Context-Free Grammar (CFG for short).
[0059] For example, in the C language, "if (condition) statement" is a structure that conforms to the syntax rules, while "ifcondition statement" is a structure that does not conform to the syntax rules. It should be noted that the methods of syntax analysis can include but are not limited to top-down (e.g., recursive descent analysis) and bottom-up (e.g., operator precedence analysis, LR analysis). In an alternative embodiment, if a Token does not conform to the syntax rules, the syntax analyzer will report a syntax error and attempt to resume analysis.
[0060] In addition to syntax analysis, in an alternative embodiment, semantic analysis of the Tokens is also required, that is, through a semantic analyzer, semantic analysis is performed on the lexical units to obtain semantic information. In a specific embodiment, the semantic analysis stage mainly checks whether the semantics of the Tokens are correct. For example, type checking, scope checking, etc. For example, in a strongly typed language, a value of string type cannot be assigned to a variable of integer type, and the semantic analyzer will detect this type mismatch error. Semantic analysis also manages the symbol table, recording the definition and usage information of identifiers. The symbol table stores information such as the name, type, and scope of identifiers for use in subsequent compilation stages.
[0061] Thus, after splitting, syntax analysis, and semantic analysis of the DSL query statement, the original DSL query statement can be split into multiple lexical units, and semantic information can be obtained. Furthermore, a syntax tree can be constructed based on the semantic information and lexical units. Furthermore, the original domain-specific query statement DSL can be parsed and reconstructed according to the astronomy knowledge graph, and the reconstructed DSL can be used for querying, avoiding the problem of writing different query statements for different data sources and improving query efficiency and accuracy.
[0062] In an alternative embodiment, constructing an astronomy knowledge graph includes: Obtaining an astronomy data set; Extracting entity fields and relationship fields from the astronomy data set; Constructing the entity fields and relationship fields into triples; Generating an astronomy knowledge graph based on the triples.
[0063] In a specific embodiment, such as Figure 2 the astronomy knowledge graph management unit 10 shown, collects a large number of astronomy data sets in the field of astronomy, which can be obtained from open-source astronomy databases or from astronomy books and other means. The application does not limit the ways to obtain astronomy data sets.
[0064] After obtaining the astronomy data set, entity fields and relationship fields are extracted from the astronomy data set. Among them, the entity fields can be terms or concepts in the field of astronomy, and the entity fields include head entity fields and tail entity fields. The relationship fields describe the relationships between the entity fields. In a specific embodiment, according to the subordinate relationship (i.e., the parent-child relationship) between the head entity and the tail entity, triples can be constructed.
[0065] For example, for the terms "right ascension", the term (MC_RA), and the term "RA", they are all a way of describing "TERM_right_ascension". Therefore, triples such as "right ascension - also known as - TERM_right_ascension" can be constructed. In this triple, it expresses that "TERM_right_ascension" has an alias of "right ascension".
[0066] In some alternative embodiments, since the knowledge graph is used to construct the astronomy concept map, and the knowledge graph data can be stored using the RDF (Resource Description Framework) data structure. RDF is a low-level data storage model. In RDF, the basic unit is a triple. The structure is "subject - predicate - object". Each triple indicates that the "subject" has an attribute "predicate", and the value of this attribute is the "object". Further, in a specific embodiment, the relationships between the terms are interconnected to complete the construction of the parent-child relationships between the terms. For example, the term "main survey telescope" is a sub-concept of "TERM_telescope". Therefore, a triple "main survey telescope - belongs to - TERM_telescope" can be constructed.
[0067] In an alternative embodiment, the graph database for the persistent storage of the astronomy knowledge graph can be a Neo4j graph database. In a specific embodiment, Cypher can be used to query the Neo4j graph database to obtain the hierarchical relationship of the astronomy domain concept graph.
[0068] It should be noted that in an alternative embodiment, in order to improve the accuracy of the knowledge graph, after the knowledge graph is initially generated, it can be verified and modified by experts.
[0069] As an alternative embodiment, parsing the archival task instruction includes: Extracting the task file from the archival task instruction; Parsing the task file to determine the access parameters of the data source device and the access logic of the file header.
[0070] In a specific embodiment, when the user has new observation data stored somewhere to be archived, the file content conforming to the YAML specification submitted by the user is received and parsed, that is, the YAML task file is extracted from the archival task instruction. Parsing the YAML task file can determine the access parameters of the data source device 2, where the access parameters may include, but are not limited to, IP address, port number, username, and password. For ease of understanding, an example will be given below.
[0071] For example, the YAML task file is as follows: ingest_job_name: "${INGEST_JOB_NAME}" source: type: oss_fits_source.OSSSource config: platform: OSS path: - s3: / / csst-prod / CSST_L0 / MSC / SCI / 60310 / 10100000000 / *.* - s3: / / csst-prod / CSST_L0 / MSC / SCI / 60310 / 10100000001 / *.* conn_config: s3_endpoint_url: "${S3_ENDPOINT}" s3_access_key_id: "${S3_ACCESS_KEY}" s3_secret_access_key: "${S3_ACCESS_SECRET}" ingest_config: tags: - CSST - fits user_props: level: "L0" path_patterns: bucket_name: "[^ / ]*\\ / ([^ / ]+)\\ / .*" sink: type: metadata-rest-api config: server: "${REST_SINK_URL}" token: "${REST_ACCESS_TOKEN}" Among them, "Platform: OSS" indicates the scanning task generated by the YAML task file, and "path" is the storage path of the file header.
[0072] In an alternative embodiment, different storage medium connection methods are different. The driver for connecting to block storage OSS can be loaded and used, and the three pieces of information, conn_config.s3_endpoint_url, conn_config.s3_access_key_id, and conn_config.s3_secret_access_key, are used to connect to block storage OSS. If "Platform: JDBC", other information not shown, such as conn_config.jdbc_url and conn_config.database_schema, will be used for connection.
[0073] In addition to the access parameters of the data source device 2, the access logic of the file header also needs to be determined. It can be understood that different file headers correspond to different parsing logics, that is, different file headers need to use different parsing logics. In a specific embodiment, the file header metadata parsing logic identifier in the content of the root YAML task file is used to load the corresponding file header metadata parsing logic. For example, the parsing logic for the FITS file header generated by this astronomical observation device. In an alternative embodiment, when the user submits the YAML task file, the management system 1's default Python parser will look for the Site-packages directory. The Site-packages directory is a special directory in the Python programming language for storing third-party modules and libraries. When the user needs to install Python packages of non-standard libraries, the files of these packages need to be placed in this directory so that the Python interpreter can find and load them. Therefore, in the above example, the YAML task file will instruct the Python interpreter to load the "OSSSource" class in the "oss_fits_source.py" file, where the OSSSource class describes how to process all the files in this scanning task.
[0074] In another alternative embodiment, when the user needs to customize the processing logic of astronomical data, Python files that implement the same interface and comply with the same specifications also need to be placed in the Site-packages folder, and the relevant folder address can be found by executing the command Python -msite. In addition, the custom file scanning logic class needs to inherit from the IngestionSourceBase class and implement various classes related to file scanning and reading under this abstract class. This application does not limit the custom methods.
[0075] It should be noted that in this example, after parsing the YAML task file, it is determined whether there is a historical alignment result for semantic alignment of the metadata in the first file header in the specified database 13. Specifically, the submitted YAML task file of the user can be used to check whether there is a semantic alignment operation result with ${INGEST_JOB_NAME} as the key-value pair key in the specified database 13 or in a service component based on an in-memory database (such as Redis).
[0076] As Figure 2 shown, in a specific embodiment, the user submits a YAML task file through the interaction unit 3, the query unit 14 parses the YAML task file, and communicates with the data heterogeneous data source device 2 according to the parsing result to scan the file headers in the data source device 2.
[0077] In the above embodiment, the management method for multi-source heterogeneous astronomical data is described in detail. This application also provides an embodiment corresponding to a management device for multi-source heterogeneous astronomical data.
[0078] Figure 4 As shown in the structural schematic diagram of a management device for multi-source heterogeneous astronomical data provided by an embodiment of this application, Figure 4 shown, the device includes: An acquisition module 40, configured to acquire an archiving task instruction for new observation data and a pre-constructed astronomy knowledge graph; A communication connection module 41, configured to parse the archiving task instruction; and establish a communication connection with a data source device for collecting new observation data according to the parsing result; A scanning module 42, configured to scan the file headers of new observation data; wherein the file headers include metadata for describing information related to new observation data; A semantic alignment module 43, configured to perform a semantic alignment operation on the metadata in the file headers based on the astronomy knowledge graph; and store the semantic alignment result in a specified database.
[0079] In addition, the management device for multi-source heterogeneous astronomical data provided by an embodiment of this application further includes: The first file header extraction module is used to extract the first file header scanned from each data source device; The first processing module is used to determine whether there is a historical alignment result for semantic alignment of the metadata in the first file header in the specified database; if it exists, the metadata in the file header is replaced according to the historical alignment result; if it does not exist, based on the astronomy knowledge graph, the specified model is used to perform semantic alignment on the metadata in the first file header; and according to the current alignment result, the other metadata in the file header except the first file header is replaced.
[0080] The alignment modification instruction acquisition module is used to obtain the user's alignment modification instruction within a preset time period; The secondary semantic alignment module is used to perform secondary semantic alignment on the metadata in the first file header according to the alignment modification instruction; and use the secondary semantic alignment result as the current alignment result; The training dataset acquisition module is used to collect the current alignment result to obtain a training dataset; The training module is used to train the specified model through the training dataset.
[0081] The query statement acquisition module is used to obtain the user's query statement for multi-source heterogeneous astronomical data; The query statement parsing module is used to parse the query statement to determine the semantic information and included word units of the query statement; The syntax tree construction module is used to construct a syntax tree for describing the hierarchical relationship between word units according to the semantic information; wherein, the syntax tree includes leaf nodes and root nodes, and the leaf nodes and root nodes are word units; The target query statement determination module is used to replace the leaf nodes with target fields in the astronomy knowledge graph to obtain a target query statement; The query result acquisition module is used to obtain the query result of multi-source heterogeneous astronomical data in the specified database through the target query statement.
[0082] The splitting module is used to split the query statement based on predefined lexical rules to obtain word units; The analysis module is used to call a syntax analyzer to perform syntax analysis on the word units and use a semantic analyzer to perform semantic analysis on the word units to obtain semantic information.
[0083] The astronomy dataset acquisition module is used to obtain an astronomy dataset; The field extraction module is used to extract entity fields and relationship fields from the astronomy dataset; The triple construction module is used to construct triples from the entity fields and relationship fields; A knowledge graph generation module for generating an astronomy knowledge graph based on triples.
[0084] A task file extraction module for extracting task files from the archived task instructions; A task file parsing module for parsing the task files to determine the access parameters of the data source device and the access logic of the file header.
[0085] Figure 5 The structural schematic diagram of a management device for multi-source heterogeneous astronomical data provided by another embodiment of the present application is as follows Figure 5 As shown, the management device for multi-source heterogeneous astronomical data includes: a memory 50 for storing computer programs; A processor 51 for implementing the steps of the management method for multi-source heterogeneous astronomical data mentioned in the above embodiment when executing the computer program.
[0086] The management device for multi-source heterogeneous astronomical data provided in this embodiment may include, but is not limited to, a laptop computer or a desktop computer, etc.
[0087] Among them, the processor 51 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 51 may be implemented in at least one hardware form of a digital signal processor (DSP for short), a field-programmable gate array (FPGA for short), and a programmable logic array (PLA for short). The processor 51 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as a central processing unit (CPU for short); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 51 may be integrated with a graphics processing unit (GPU for short), and the GPU is responsible for the rendering and drawing of the content to be displayed on the display screen. In some embodiments, the processor 51 may also include an artificial intelligence (AI for short) processor, and the AI processor is used to process the computing operations related to machine learning.
[0088] The memory 50 may include one or more computer-readable storage media, which may be non-transitory. The memory 50 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 50 is at least used to store the following computer program 501. After the computer program is loaded and executed by the processor 51, it can implement the relevant steps of the multi-source heterogeneous astronomical data management method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 50 may also include an operating system 502, data 503, etc., and the storage method may be transient storage or permanent storage. Among them, the operating system 502 may include Windows, Unix, Linux, etc. The data 503 may include, but is not limited to, the relevant data involved in the multi-source heterogeneous astronomical data management method, etc.
[0089] In some embodiments, the multi-source heterogeneous astronomical data management device may further include a display screen 52, an input / output interface 53, a communication interface 54, a power supply 55, and a communication bus 56.
[0090] Those skilled in the art can understand that Figure 5 the structure shown in does not constitute a limitation on the multi-source heterogeneous astronomical data management device, and may include more or fewer components than shown in the figure.
[0091] The multi-source heterogeneous astronomical data management device provided by the embodiments of the present application includes a memory and a processor. When the processor executes the program stored in the memory, it can implement the multi-source heterogeneous astronomical data management method in the above embodiments.
[0092] It should be noted that although the operations are depicted in a specific order in the drawings, this should not be construed as requiring the operations to be performed in the specific order shown or sequentially, or requiring all of the illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of the various system modules and components in the above embodiments should not be understood as required in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
Claims
1. A management method for multi-source heterogeneous astronomical data, characterized in that, The method includes: Obtaining an archiving task instruction for new observation data and a pre-constructed astronomy knowledge graph; Parsing the archiving task instruction; and establishing a communication connection with the data source device that collects the new observation data according to the parsing result; Scanning the file header of the new observation data; wherein, the file header includes metadata for describing information related to the new observation data; Based on the astronomy knowledge graph, performing a semantic alignment operation on the metadata in the file header; and storing the semantic alignment result in a specified database.
2. The management method of multi-source heterogeneous astronomical data according to claim 1, characterized in that, Performing a semantic alignment operation on the metadata in the file header based on the astronomy knowledge graph, including: Extracting the first file header scanned from each of the data source devices; Determining whether there is a historical alignment result in the specified database for performing a semantic alignment operation on the metadata in the first file header; If so, replacing the metadata in the file header according to the historical alignment result; If not, based on the astronomy knowledge graph, performing a semantic alignment operation on the metadata in the first file header through a specified model; and replacing the metadata in the file header other than the first file header according to the current alignment result.
3. The management method of multi-source heterogeneous astronomical data according to claim 2, characterized in that, Before replacing the metadata in the file header other than the first file header according to the current alignment result, it further includes: If an alignment modification instruction from the user is obtained within a preset time period; Performing a secondary semantic alignment on the metadata in the first file header according to the alignment modification instruction; and using the secondary semantic alignment result as the current alignment result; Collecting the current alignment result to obtain a training data set; Training the specified model through the training data set.
4. The management method of multi-source heterogeneous astronomical data according to claim 1, wherein, The method further includes: Obtaining a query statement of the user for multi-source heterogeneous astronomical data; Parsing the query statement to determine the semantic information and included word units of the query statement; According to the semantic information, constructing a syntax tree for describing the hierarchical relationship between the word units; wherein, the syntax tree includes leaf nodes and a root node, and the leaf nodes and the root node are the word units; Replacing the leaf nodes with target fields in the astronomy knowledge graph to obtain a target query statement; Obtaining a query result of multi-source heterogeneous astronomical data in the specified database through the target query statement.
5. The management method of multi-source heterogeneous astronomical data according to claim 4, wherein, Parsing the query statement to determine the semantic information and included word units of the query statement, including: Based on predefined lexical rules, splitting the query statement to obtain the word units; Invoking a syntax analyzer to perform syntax analysis on the word units, and through a semantic analyzer, performing semantic analysis on the word units to obtain the semantic information.
6. The management method of multi-source heterogeneous astronomical data according to claim 1, characterized in that, Constructing an astronomy knowledge graph, including: Obtaining an astronomy data set; Extracting entity fields and relationship fields from the astronomy data set; Constructing the entity fields and the relationship fields into triples; Generating an astronomy knowledge graph based on the triples.
7. The management method of multi-source heterogeneous astronomical data according to claim 1, characterized in that, Parsing the archiving task instruction, including: Extracting the task file in the archiving task instruction; Parse the task file to determine the access parameters of the data source device and the access logic of the file header.
8. A management device for multi-source heterogeneous astronomical data, characterized in that The device includes: An acquisition module, configured to acquire an archiving task instruction for new observation data and a pre-constructed astronomical knowledge graph; A communication connection module, configured to parse the archiving task instruction; and establish a communication connection with the data source device that collects the new observation data according to the parsing result; A scanning module, configured to scan the file header of the new observation data; wherein the file header includes metadata for describing information related to the new observation data; A semantic alignment module, configured to perform a semantic alignment operation on the metadata in the file header based on the astronomical knowledge graph; and store the semantic alignment result in a specified database.
9. A management device for multi-source heterogeneous astronomical data, comprising a memory and a processor, where a computer program that can run on the processor is stored on the memory, characterized in that, When the processor executes the program, it implements the steps of the management method for multi-source heterogeneous astronomical data according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the management method for multi-source heterogeneous astronomical data according to any one of claims 1 to 7.
Citation Information
Patent Citations
Autonomous data lake construction system and method based on associated data
CN110941612A
Knowledge graph construction method and system based on multi-source database
CN114201616A
Lake and warehouse integrated multi-source remote sensing space-time big data processing method and device
CN116303249A
Data processing method and device, equipment and medium
CN116483850A
Digital base fusion system and electronic equipment
CN119760007A
Cited By
Astronomical knowledge graph construction method and device, astronomical knowledge graph retrieval method and device and medium
CN121351964A