A digital resource collection method and system for science museums
By using the JSON-LD data format and intelligent processing system, the problems of low efficiency and lack of flexibility in the acquisition of digital museum resources have been solved. It has achieved unified access and resource association of multi-source heterogeneous data, built an interconnected digital cultural ecosystem, and ensured the long-term reliability of metadata and the traceability of resources.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2026-04-07
AI Technical Summary
Existing digital museums suffer from inefficiency, error-proneness, and lack of flexibility and intelligence in resource acquisition and metadata construction. They struggle to handle diverse and complex digital resources, especially unstructured data, and the metadata construction process is difficult to personalize.
The system uses the JSON-LD data format to collect and process multi-source heterogeneous data. It connects to multiple data sources through a protocol adaptation layer, generates standardized basic and original feature metadata, generates type labels by combining a rule engine and a classification model, extracts structured metadata, and generates supplementary metadata through a multimodal parsing pipeline to achieve resource association.
It enables intelligent processing of digital resources across institutions, breaks down data silos, provides a digital cultural ecosystem that maximizes value through cross-temporal and spatial sharing, and ensures the long-term reliability of metadata and the traceability of resources.
Smart Images

Figure CN120951049B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the technical field of digital museums and digital exhibition, education, collection, research, and cultural and creative services, and in particular to a digital resource collection technology for scientific museums. BACKGROUND
[0002] In the current construction of digital museums, resource collection and metadata construction technologies mainly rely on standardized metadata frameworks (such as Dublin Core and CIDOC CRM) and traditional manual annotation methods. These methods can provide basic resource descriptions, but have obvious shortcomings when dealing with diversified and complex digital resources. For example, the processing of unstructured data (such as pictures, videos, and audio) still relies on manual analysis, which is inefficient and prone to errors. Moreover, existing systems often lack flexible scalability and intelligence, making it difficult to adapt to new resource types or rapidly changing needs.
[0003] In addition, the metadata construction process of existing technologies is relatively fixed, and it is difficult to automatically generate personalized fields according to the specific characteristics of resources, resulting in insufficient support for complex and multi-modal resources. SUMMARY
[0004] The purpose of embodiments of the present disclosure is to provide a digital resource collection method and system for scientific museums. The digital resource collection of scientific museums is not simply "data collection", but rather to transform scattered resources into a knowledge network that can be associated.
[0005] According to one aspect of the present disclosure, a digital resource collection method for scientific museums is provided, wherein the method comprises the following steps:
[0006] Accessing multi-source heterogeneous data and generating first JSON-LD data for each accessed data, including standardized basic metadata and original feature metadata;
[0007] Based on the first JSON-LD data, extracting explicit technical features from the original feature metadata to generate first type labels, and parsing semantic information in the standardized basic metadata to generate second type labels, fusing the first type labels and the second type labels to obtain final type labels, forming second JSON-LD data;
[0008] Based on the second JSON-LD data, loading a predefined rule template according to the type labels to extract structured metadata, including type core metadata and type specific metadata, and retaining references to original access data, forming third JSON-LD data;
[0009] Based on the third JSON-LD data, the corresponding pipeline branch is determined from the multimodal parsing pipeline according to the type label. The original access data is parsed according to the structured metadata using the pipeline branch to generate supplementary metadata and form the fourth JSON-LD data.
[0010] Based on the fourth JSON-LD data of each access resource, a correlation is established between the metadata of different access resources to form the fifth JSON-LD data of each access resource.
[0011] According to one aspect of this disclosure, a resource acquisition system for a digital museum is also provided, wherein the system includes a memory and a processor, the memory storing computer program instructions, and when the computer program instructions are executed by the processor, the system is configured to perform the following operations:
[0012] Access multi-source heterogeneous data and generate first JSON-LD data for each accessed data, including standardized basic metadata and original feature metadata;
[0013] Based on the first JSON-LD data, explicit technical features are extracted from the original feature metadata to generate a first type label, and semantic information in the standardized basic metadata is parsed to generate a second type label. The first type label and the second type label are then merged to obtain the final type label, forming the second JSON-LD data.
[0014] Based on the second JSON-LD data, a predefined rule template is loaded according to the type tag to extract structured metadata, including type core metadata and type-specific metadata, while retaining the reference to the original access data to form the third JSON-LD data;
[0015] Based on the third JSON-LD data, the corresponding pipeline branch is determined from the multimodal parsing pipeline according to the type label. The original access data is parsed according to the structured metadata using the pipeline branch to generate supplementary metadata and form the fourth JSON-LD data.
[0016] Based on the fourth JSON-LD data of each access resource, a correlation is established between the metadata of different access resources to form the fifth JSON-LD data of each access resource.
[0017] The embodiments disclosed herein construct an intelligent processing system that spans the entire data lifecycle in the integration of digital resources across institutions, achieving a leap from heterogeneous access to resource association. By uniformly converting various heterogeneous resources from planetariums, natural history museums, and science museums into structured semantic data (JSON-LD data), the digital resource acquisition system breaks through the limitations of traditional data silos. The core value of digital museums lies in breaking down physical space limitations, and the unified JSON-LD semantic data format provides a "common language" for cross-museum resource collaboration.
[0018] Digital museums need to address long-term challenges such as resource format iteration and storage media aging. However, the structured semantic characteristics of JSON-LD naturally possess "format independence." Even if the original resource files (such as early 3D model formats) cannot be directly read due to technological obsolescence, the metadata they contain, such as "creation date," "scientific background," and "related entities," can still be preserved for a long time and parsed by new systems.
[0019] These technological effects have collectively propelled digital museums from "a collection of scattered digital exhibits" to "an interconnected and dynamically growing digital cultural ecosystem," ultimately enabling cross-temporal and spatial sharing of cultural heritage resources and maximizing their value. Attached Figure Description
[0020] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0021] Figure 1 A flowchart illustrating a digital resource acquisition method for a science museum according to one embodiment of the present disclosure is shown.
[0022] The same or similar reference numerals in the accompanying drawings represent the same or similar parts. Detailed Implementation
[0023] The specific embodiments of this disclosure will be further described below with reference to the accompanying drawings.
[0024] Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments of this disclosure are described as apparatuses represented by block diagrams and processes or methods represented by flowcharts. Although the flowcharts depict the operation processes of the various embodiments of this disclosure as sequential processes, many of the operations may be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations may be rearranged. The processes of the various embodiments of this disclosure may be terminated when their operations are completed, but may also include additional steps not shown in the flowcharts. The processes of the various embodiments of this disclosure may correspond to methods, functions, procedures, subroutines, subroutines, etc.
[0025] The methods illustrated in the flowcharts and the apparatuses illustrated in the block diagrams discussed below can be implemented in hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, program code or code segments that perform the necessary tasks can be stored in a machine or a computer-readable medium such as a storage medium. One or more processors can perform the necessary tasks.
[0026] Similarly, it will also understand any flowchart, state transition diagram, and the like, representing various processes that can be adequately described as program code stored in a computer-readable medium and thus executed by a computer device or processor, whether or not such computer device or processor is explicitly shown.
[0027] In this document, the term "storage medium" can refer to one or more devices for storing data, including read-only memory (ROM), random access memory (RAM), magnetic RAM, core memory, disk storage media, optical storage media, flash memory devices, and / or other machine-readable media for storing information. The term "computer-readable medium" may include, but is not limited to, portable or fixed storage devices, optical storage devices, and various other media capable of storing and / or containing instructions and / or data.
[0028] A code segment can represent a procedure, function, subroutine, program, routine, subroutine, module, software package, class, or any combination of instructions, data structures, or program descriptions. A code segment can be coupled to another code segment or hardware circuit by passing and / or receiving information, data, arguments, parameters, or stored content. Information, arguments, parameters, data, etc., can be passed, forwarded, or transmitted via any suitable means, including storage sharing, message passing, token passing, network transmission, etc.
[0029] In this context, "computer device" refers to an electronic device that can perform predetermined processing procedures such as numerical calculations and / or logical calculations by running predetermined programs or instructions. It may include at least a processor and a memory, wherein the predetermined processing procedures are performed by the processor executing program instructions pre-stored in the memory, or by hardware such as ASIC, FPGA, DSP, etc., or by a combination of the above.
[0030] The term "computer device" as used above is generally embodied in the form of a general-purpose computer device, whose components may include, but are not limited to, one or more processors or processing units and system memory. System memory may include computer-readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. The "computer device" may further include other removable / non-removable, volatile / non-volatile computer-readable storage media. The memory may include at least one computer program product having a set (e.g., at least one) of program modules configured to perform the functions and / or methods of the embodiments of this disclosure. The processor executes various functional applications and data processing by running programs stored in the memory.
[0031] For example, a computer program for performing various functions and processes of multiple embodiments of the present disclosure is stored in the memory, and when the processor executes the corresponding computer program, the digital resource acquisition system of the present disclosure is implemented.
[0032] Typically, computer devices can be user devices or network devices, or even a combination of both. User devices include, but are not limited to, personal computers (PCs), laptops, and mobile terminals; mobile terminals include, but are not limited to, smartphones and tablets. Network devices include, but are not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing, which is a type of distributed computing consisting of a super virtual computer composed of a group of loosely coupled computers. The computer devices can operate independently to implement the embodiments of this disclosure, or they can connect to a network and implement the embodiments of this disclosure through interaction with other computer devices in the network. The networks in which the computer devices reside include, but are not limited to, the Internet, wide area networks (WANs), metropolitan area networks (MANs), local area networks (LANs), and VPN networks.
[0033] It should be noted that the user equipment, network equipment, and network mentioned are merely examples. Other existing or future computing devices or networks that are applicable to the embodiments of this disclosure should also be included within the scope of protection of this disclosure and are incorporated herein by reference.
[0034] The specific structural and functional details disclosed herein are merely representative and are intended to describe exemplary embodiments of this disclosure. However, the various embodiments of this disclosure can be implemented in many alternative forms and should not be construed as being limited solely to the embodiments set forth herein.
[0035] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0036] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. Unless the context clearly indicates otherwise, the singular forms “a” and “an” as used herein are also intended to include the plural. It should also be understood that the terms “comprising” and / or “including” as used herein specify the presence of the stated features, integers, steps, operations, units, and / or components, without excluding the presence or addition of one or more other features, integers, steps, operations, units, components, and / or combinations thereof.
[0037] It should also be mentioned that in some alternative implementations, the functions / actions mentioned may occur in a different order than those shown in the figures. For example, depending on the functions / actions involved, the two figures shown successively may actually be executed substantially simultaneously or sometimes in reverse order.
[0038] See Figure 1 The illustration shows a digital resource acquisition process for science museums according to an embodiment of the present disclosure.
[0039] like Figure 1As shown, in step S1, the digital resource acquisition system adapts to access multi-source heterogeneous data through protocol adaptation and generates first JSON-LD data for each accessed data, including standardized basic metadata and original feature metadata; in step S2, based on the first JSON-LD data, the digital resource acquisition system extracts explicit technical features from the original feature metadata to generate a first type label, and parses the semantic information in the standardized basic metadata to generate a second type label. The first type label and the second type label are then merged to obtain the final type label, forming second JSON-LD data; in step S3, based on the second JSON-LD data, the digital resource acquisition system loads predefined rule modules according to the type labels. The system extracts structured metadata, including type core metadata and type-specific metadata, and retains references to the original access data to form third JSON-LD data. In step S4, the digital resource acquisition system, based on the third JSON-LD data, determines the corresponding pipeline branch from the multimodal parsing pipeline according to the type tag, and uses the pipeline branch to parse the original access data according to the structured metadata to generate supplementary metadata, forming fourth JSON-LD data. In step S5, the digital resource acquisition system, based on the fourth JSON-LD data of each access resource, establishes associations between the metadata of different access resources to form fifth JSON-LD data of each access resource.
[0040] Specifically, in step S1, the digital resource acquisition system adapts to access multi-source heterogeneous data through protocol and generates first JSON-LD data for each accessed data, including standardized basic metadata and original feature metadata.
[0041] This publicly available digital museum is not based on traditional museums with cultural relics collections, but rather on a comprehensive information system / platform for natural science museums such as science and technology museums, natural history museums, and planetariums. It uses intelligent technology to systematically collect, integrate, and manage multimodal and heterogeneous digital resources.
[0042] Based on the actual needs of natural science museums, the types of resources typically include collections, exhibitions, educational activities, academic documents, multimedia resources, cultural and creative products, and social media content.
[0043] Resources come from various systems and platforms, and their data storage formats differ. For example, a museum's collection management system may use a relational database (such as MySQL), while multimedia resources typically use unstructured data formats (such as JPEG, MP4, WAV), or even PDFs and scanned documents. This heterogeneity necessitates addressing issues such as data format conversion, information integration, and normalization during data acquisition.
[0044] Inconsistent standards among data sources are also a major challenge. For example, different venues may use different classification systems to describe the same type of resource, resulting in even digital resources of the same type being stored with different field names and structures.
[0045] In the data integration architecture of natural science museums, the core task of step S1 is to build a unified access layer for multi-source heterogeneous data, and generate basic data carriers with standardized features and original traceability capabilities through protocol adaptation and semantic mapping. This process needs to address the complex characteristics of multi-dimensional data sources both inside and outside the museum, and form a scalable metadata governance framework.
[0046] To achieve unified access to multi-source digital resources, the protocol adaptation layer shields the differences in underlying data sources through a unified interface. These data sources include structured business systems within the venue, API-type data sources from external platforms, and various unstructured multimedia data streams.
[0047] For structured data sources, the core characteristic is that the data itself has a clear table structure and fixed schema constraints. Typical structured data sources include museum collection management systems and exhibition management systems. The underlying storage of these management systems is usually relational databases (such as MySQL and PostgreSQL). For these data sources, the protocol adaptation layer prioritizes using JDBC / ODBC connectors to directly establish long-lived database connections and extracts data periodically or in batches as needed using predefined SQL templates. During the data extraction process, to ensure the traceability and consistent association of data across systems, the protocol adaptation layer maps key fields of each record (such as collection_id) to globally standardized unique identifiers and completely preserves the original records in the _original node, ensuring that the original structure and data content are not lost.
[0048] API-based data sources refer to data sources that rely entirely on external open interfaces. These data sources are not directly linked to the museum's management system and lack a local storage system. Typical examples include museum websites, open data platforms, and various social media platforms (such as WeChat and Weibo). The protocol adaptation layer's access strategy for API-based data sources emphasizes obtaining dynamic access tokens through authentication methods such as OAuth 2.0 and calling the platform's standard RESTful API or SOAP interfaces to retrieve data. For the response results, the protocol adaptation layer automatically converts the data returned by the API (JSON, XML, Protobuf, etc.) to a standardized JSON format and designs an incremental polling mechanism for pagination interfaces, combining request parameters and response timestamps to form a complete time-series record.
[0049] For unstructured data streams, these data sources primarily consist of multimedia data such as audio and video, typically including scanned images and scanned copies of historical archives during the digitization process. The access strategy for this type of data source emphasizes standardized processing of the original multimedia files and metadata generation. The protocol adaptation layer uses multimedia tools such as FFmpeg to extract keyframes and audio clips in real-time or on demand, and extracts key technical metadata such as resolution, bitrate, and sampling rate. The original binary files are Base64 encoded and stored along with the metadata in an intermediate database or object storage system. Furthermore, for unstructured data such as scanned images and PDF documents, the protocol adaptation layer further supports batch import and content parsing, combining OCR technology, image processing technology, and PDF parsers to convert text and image content into structured text or semi-structured document description information.
[0050] Regardless of whether the data source is structured, API-based, or unstructured, the protocol adaptation layer performs a unified data cleaning and standardization process, including data format conversion, field standardization, data deduplication, and outlier detection. This ensures that data from different sources and in different formats can be transformed into standardized datasets with a unified structure and complete content, providing high-quality foundational data support for subsequent metadata extraction, resource management, and business analysis. Ultimately, the protocol adaptation layer not only shields the diverse differences of the underlying data sources but also provides consistent, traceable, and scalable data input capabilities to the upper-layer digital resource management platform through standardized data models and unified access processes.
[0051] The first JSON-LD data generated for each access data is ultimately divided into two layers:
[0052] 1) Standardized basic metadata consists of common fields for all resource types, containing only business semantic fields common to all resources, and is independent of specific technical implementations.
[0053] According to one example, standardized basic metadata may include a globally unique identifier (id), a resource title (title), a resource description (description), and a resource creation time (created, which must be forcibly converted to ISO 8601 format).
[0054] 2) Original feature metadata records the technical characteristics and access details of the resource, and is stored according to data source type:
[0055] The original characteristic metadata fields and their descriptions for file-type resources are exemplified below:
[0056] Fields describe path File storage path (e.g., / planetarium / 2023_mars_show / 4k_video.mp4) extension File extension (e.g., .mp4) mime_type Media type (e.g., video / mp4) file_size File size (bytes) hash Hash value of file content (e.g., sha256:abc123) exif Technical metadata (e.g., {"resolution": "3840x2160", "frame_rate": "30fps"})
[0057] The original characteristic metadata fields and their descriptions for API-type resources are exemplified below:
[0058] Fields describe endpoint API request path (e.g., / cgi-bin / material / get) http_method Request method (e.g., GET / POST) request_headers Request headers (e.g., {"Authorization": "Bearer token123"}) response_status Response status code (e.g., 200) response_body Raw response body (unconverted data such as JSON / XML)
[0059] Examples of the original characteristic metadata fields and their descriptions for database-type resources are as follows:
[0060] Fields describe table_name Source table name (e.g., collection_info) schema Table structure definition (field names, types, constraints) query SQL statement for data extraction (e.g., SELECT name FROM specimen_collection WHERE location='Shanghai Natural History Museum') transaction_id Database transaction ID (used for tracing)
[0061] In the resource acquisition process of digital museums, protocol adaptation and hierarchical metadata design provide key technical support for the integration of heterogeneous data from multiple sources, such as science and technology museums, natural history museums, and planetariums. By adapting to different protocols (such as HTTPAPI, database connections, and file transfer protocols), the heterogeneous resources of these natural science museums are uniformly accessed, solving the "data silo" problem caused by format differences in traditional multi-source data. For example, JSON-formatted observation data provided by planetariums and XML specimen archives from natural history museums can be transformed into a unified structured expression through the protocol parsing layer, laying the foundation for subsequent processing.
[0062] This step employs a layered metadata architecture in JSON-LD format, decoupling standardized basic metadata from raw feature metadata. The standardization layer focuses on core business semantics across domains, such as globally unique identifiers, resource titles, multilingual descriptions, and ISO standard timestamps—common fields that ensure comparability of resources from different venues across fundamental dimensions. For example, meteorite sample images from the planetarium and electromagnetic experiment videos from the science museum both include a "creation time" field in the standardization layer and are forcibly converted to ISO 8601 format, enabling subsequent time-based resource correlation analysis. This design preserves the purity of business semantics while avoiding interference from different technical implementation details on core data.
[0063] The original feature metadata layer fully records the technical fingerprints of various resources. For example, high-resolution specimen images from natural history museums retain the EXIF information of the imaging equipment, while radio telescope data from planetariums stores observation frequency bands and sampling accuracy parameters.
[0064] This hierarchical storage strategy not only meets the rigid requirement for standardization of core metadata when integrating resources across institutions, but also provides a complete information foundation for technology tracing in specific scenarios. When it is necessary to reproduce certain astronomical observation data, the telescope parameters stored in the original feature layer can be directly used for data verification without affecting the common business fields in the standardization layer.
[0065] Subsequently, in step S2, the digital resource acquisition system extracts explicit technical features from the original feature metadata based on the first JSON-LD data to generate a first type label, and parses the semantic information in the standardized basic metadata to generate a second type label. The first type label and the second type label are then merged to obtain the final type label, forming the second JSON-LD data.
[0066] Based on the first JSON-LD data generated in step S1, the system completes the accurate identification and correction of resource types through the collaborative processing of the rule engine and the classification model.
[0067] First, the system generates the first type of label by matching explicit feature trigger rules in the original feature metadata.
[0068] The design goal of the rule engine is to extract explicit technical features from raw feature metadata and quickly generate first-type labels (i.e., initial type labels). The rule base is predefined according to resource type and uses technologies such as regular expressions, field matching, and path pattern recognition.
[0069] Extension matching: When _original.file.extension is mp4, the multimedia resource class tag is triggered; if the path contains / exhibits / space / blackhole_simulation.mp4, the science video subclass is overlaid.
[0070] API endpoint mapping: If _original.api.endpoint matches / cgi-bin / exhibition / , it is marked as an exhibition class, and further refined into a virtual exhibition or a physical exhibition based on whether the virtual_link field exists in the response body.
[0071] Database table name inference: Check if _original.database.table_name contains specimen_collection, trigger collection category tags, and combine with the geological_era field value (such as "Jurassic") to improve the confidence of paleontological collections.
[0072] For example, for a tweet data accessed from the WeChat API, the original metadata has `_original.api.endpoint` as ` / cgi-bin / material / get`. The rule engine generates an initial type tag based on the pre-configured interface path rules: Social Media_WeChat Official Account. If the data's `_original.file.extension` is `mp4` (attached video), then a video category tag is appended, forming a composite initial type.
[0073] Furthermore, a second type of label is obtained by parsing the semantic information in the standardized basic metadata through a classification model.
[0074] The goal of classification models is to deduce resource types from the semantic information of standardized basic metadata. Training strategies for classification models include, for example, using a large amount of museum literature (such as science museum lab reports and natural history museum specimen records) for domain-adaptive pre-training on top of the general BERT model, enabling the model to understand domain terms such as "qubit manipulation" and "fossil tomography."
[0075] The final type label is obtained by fusing the first type label and the second type label. The fusion strategy needs to balance efficiency and accuracy.
[0076] Weighting: The system assigns 60% confidence weight to the rule engine, prioritizing the reliability of strong feature signals such as file extensions and path patterns, while reserving 40% weight for the semantic understanding results of the classification model, balancing the decision-making basis of explicit rules and implicit semantics. The final type label is determined based on the weighted confidence scores of the first and second type labels.
[0077] Threshold coverage: When the confidence level of the second type label exceeds a predetermined threshold, the second type label is used as the final type label. If the classification model's confidence level for a certain type exceeds 0.9, the rule result is directly overridden to avoid misclassification of explicit features.
[0078] Hierarchical inheritance: The initial labels output by the rule engine (first-type labels) serve as the parent class, and the classification model results (second-type labels) serve as the child classes. These are concatenated to form hierarchical type labels, such as a parent-child structure. For example, the video class generated by the rules and the science education output by the classification model are merged into a single video class_science education.
[0079] For example, in JSON data from the Weibo platform, `_original.api.endpoint` matches ` / weibo / museum / `, triggering the social media content category tag in the rules engine. The classification model extracts the science popularization live stream semantics from the tweet "#ScienceMuseumLive# Today we reveal the working principle of the 'quantum computing prototype'!", and after merging, generates a composite tag: social media content_science popularization live stream.
[0080] For example, an H5 interactive courseware with the path / education / astronomy / generates a multimedia resource named "Astronomy Education" based on the path using a rule engine. The classification model identifies STEM education features from the description "Simulation of Planetary Orbit Dynamics in the Solar System" and outputs "Educational Activity - Astrophysics". Because the model confidence score of 0.92 exceeds the threshold, the final type is corrected to "Educational Activity - Astrophysics", while retaining the multimedia resource tag for format retrieval.
[0081] The final type label is written to the @type field, and the source of the decision can be recorded, thus obtaining the second JSON-LD data.
[0082] Based on an example, type tags can be divided into three levels:
[0083] Primary type: Document class (the basic type for rule triggering).
[0084] Secondary types: Papers, Reports, Popular Science (subclasses of model recognition).
[0085] Level 3 types: research reports, journal articles, and technical white papers (fine-grained tags required for business scenarios).
[0086] For example,
[0087] 1) Rule engine processing (driven by raw feature metadata)
[0088] Rule 1: .pdf extension → triggers document class tags.
[0089] Rule 2: Database table name research_papers → Overlay academic attributes.
[0090] First type of tag: Documents_Academic (coarse granular).
[0091] 2) Classification model processing (driven by standardized metadata)
[0092] Input features:
[0093] Title keywords: spectroscopy, alloy composition → pointing to materials science research.
[0094] Description text: X-ray fluorescence spectrometry, component detection → Identifying experimental analysis papers.
[0095] Model output:
[0096] Second type of tag: Academic paper: 0.91 (fine-grained subclass: research report).
[0097] 3) Tag Fusion (Decision Logic and Type Correction)
[0098] Rules and model weight allocation:
[0099] Rule confidence weight: 0.5 (explicit technical features are reliable but coarse-grained).
[0100] Model confidence weight: 0.5 (accurate semantic understanding but dependent on data quality).
[0101] Conflict resolution strategies:
[0102] If the model's confidence level for fine-grained types exceeds a threshold (e.g., ≥0.85), then the coarse-grained results of the rules are overridden.
[0103] Final output:
[0104] Because the score of Academic Paper (0.91) exceeds the threshold, the final type is Paper_Research Report, and it inherits the academic attribute of the rule tag.
[0105] When the model's confidence level for fine-grained types is significantly higher than that of the rule results (e.g., 0.91 > threshold 0.85), the model output is used directly to avoid the loss of information from coarse-grained labels.
[0106] Hierarchical type tags can simultaneously retain the contributions of rules and models, allowing downstream processes to use them as needed (e.g., filtering by document type during retrieval, and focusing on research reports during analysis).
[0107] The final tag format is a composite structure as follows:
[0108] {
[0109] "@type": ["Document type", "Paper type", "Research report"],
[0110] "mu:type_source": {
[0111] "rule_engine": ["document class"],
[0112] "classification_model": ["paper type_research report"]
[0113] }
[0114] }
[0115] The dual-engine type annotation mechanism effectively solves the problem that traditional single-classification methods struggle to handle the semantic complexity of specialized resources. The rule engine quickly anchors the basic resource type by parsing original technical features. This rapid localization mechanism based on explicit features provides efficient technical support for the preliminary classification of massive heterogeneous resources. The domain classification model overcomes the limitations of general natural language processing through domain knowledge injection, demonstrating accurate semantic parsing capabilities within the specialized context of natural science museums.
[0116] The hierarchical design of the fusion strategy balances the dual requirements of efficiency and accuracy, forming a classification system that is both stable and adaptable. The weight allocation mechanism prioritizes decision weights with strong technical features, while the threshold coverage rule activates error correction when the semantic confidence is extremely high. The hierarchical inheritance structure achieves fine-grained management through nested parent and child classes.
[0117] In step S3, the digital resource acquisition system extracts structured metadata based on the second JSON-LD data and loads predefined rule templates according to type tags. This includes type core metadata and type-specific metadata, while retaining references to the original access data, thus forming the third JSON-LD data.
[0118] Here, based on the second JSON-LD data output in step S2, the system uses type tags to drive the precise extraction of structured metadata. This step extracts explicit structured information from the original resource data through predefined rules, providing a standardized and programmable business semantic framework for subsequent processing while maintaining the integrity and traceability of the original resource data.
[0119] Structured information extraction based on explicit rules relies on deterministic methods such as field path matching and regular expressions. Specifically, YAML rule templates can be used to extract structured metadata based on type tags.
[0120] The essence of YAML rule templates is a type-driven metadata extraction blueprint. Each resource type (such as collections, exhibitions, and educational activities) corresponds to an independent template file, defining the core metadata that must be extracted and the optional type-specific metadata for that type. Templates describe field mapping rules, data cleaning logic, and validation conditions using declarative syntax, and their design follows three main principles:
[0121] 1) Domain adaptability: The field definitions are tailored to the various business needs of natural science museums.
[0122] 2) Extraction accuracy: JSONPath / XPath is used to accurately locate the target field in the original data, avoiding noise caused by fuzzy matching.
[0123] 3) Scalability: Supports dynamic loading of new type templates to adapt to future additions of resource types (such as the "digital twin model" class that may be added in the future).
[0124] Based on the `@type` main tag in the second JSON-LD (e.g., "Collection Category"), load the corresponding YAML rule file for that type. If the resource has multiple tags (e.g., "Video Category_Science Popularization"), select the main type template according to priority.
[0125] Regarding the extraction of core metadata for types, based on an example, for fields defined in `core_metadata`, the raw values are extracted from the original data according to the source path. For example, based on the resource "Rocket Model Test Video" from the Science and Technology Museum, whose resource type is "Multimedia Resource", the raw value "00:05:40" is extracted according to the `duration` field.
[0126] It can also perform data cleaning. For example, it can use regular expressions for text matching or validate data using a validator. For instance, it can correct "resolution: 1920×1080" in video parameters to the standard term "1080p (FHD)".
[0127] Regarding the extraction of type-specific metadata, based on an example, the `type_specific_metadata` rule is parsed to extract extended attributes from the raw data. For example, the raw value 29.97 is extracted based on the `frame_rate` field.
[0128] Retaining references to the original resource data can be achieved through a data tiered storage strategy.
[0129] The original feature metadata from step S1 is completely copied to the _original node of the third JSON-LD, including the initial access data without any transformation. For example, the raw JSON body of the API response, or the binary data stream of the file.
[0130] Taking exhibition resources as an example, its YAML template definition is as follows:
[0131] Type core metadata: exhibition period, physical exhibition hall number, curator.
[0132] Type-specific metadata: Virtual exhibition hall links, visitor statistics, list of partner organizations.
[0133] When handling an online exhibition accessed via the WeChat API:
[0134] Based on @type: Exhibition_Virtual_Load Template, extract start_date and end_date from the API response body and merge them into ISO 8601 time period format.
[0135] Extract the names of partner museums from the partner_list field of _original.api.response.body, and filter out unauthorized institutions.
[0136] The original API response is fully preserved in the _original node for future expansion.
[0137] The third JSON-LD data generated in step S3 has dual characteristics:
[0138] Business-ready structured data: Type core metadata and type-specific metadata can directly drive front-end applications, such as filtering multimedia resources based on the resolution field.
[0139] Original data traceability: The _original node supports data quality backtracking and reprocessing. For example, when the resolution cleaning rules change, the data can be recalculated based on the original values.
[0140] The type-driven, structured metadata extraction mechanism effectively solves the challenge of describing the multi-dimensional characteristics of professional resources. Through declarative configuration of YAML rule templates, the system can automatically adapt differentiated metadata extraction strategies for different types of resources, meeting the basic needs of cross-institutional data comparison while preserving the unique characteristics of various resources. This type-based processing architecture significantly improves the efficiency of metadata governance for complex resource systems. Simultaneously, the layered data storage strategy constructs a complete digital resource traceability chain by retaining original resource references.
[0141] Next, in step S4, the digital resource acquisition system, based on the third JSON-LD data, determines the corresponding pipeline branch from the multimodal parsing pipeline according to the type label, and uses the determined pipeline branch to parse the original access data according to the structured metadata extracted in step S3, generating supplementary metadata to form the fourth JSON-LD data.
[0142] Here, based on the third JSON-LD data generated in step S3, the system selects and executes a multimodal parsing pipeline driven by type tags, and combines structured metadata to parse deep semantic features from the original data content, thereby expanding the computability and relevance of resources. Structured metadata provides domain-prior knowledge and algorithmic constraints, such as those for optimizing algorithm parameters.
[0143] As an example, the selection of a 3D model analysis pipeline is determined by type tags (such as "dinosaur fossils"); the era field (such as "Cretaceous") constrains the parameter range of the material analysis algorithm. Structured metadata serves as context, optimizing algorithm parameters. For instance, if the exhibition duration field includes holidays, the visitor flow prediction model will incorporate a holiday factor.
[0144] The multimodal parsing pipeline is a processing framework driven by both type and data form. Each pipeline branch is designed for a specific combination of resource type (such as collections, exhibitions) and data modality (such as video, 3D models, text).
[0145] The multimodal parsing pipeline is a modular processing system designed for various resource types in science museums, including collections, exhibitions, educational activities, academic documents, multimedia resources, cultural and creative products, and social media content. It includes core branches such as text parsing, multimedia parsing, and 3D model parsing, with each branch precisely bound to the resource type through type tags.
[0146] For example, a resource of type "Collection_Dinosaur Fossil_3D Model" may trigger a 3D model parsing pipeline: extracting features such as geometric shape and bone surface texture.
[0147] Some examples of pipeline branches are shown below:
[0148] Video / audio parsing pipeline
[0149] Input: Raw video file (MP4), structured metadata (resolution, duration).
[0150] Processing steps:
[0151] Keyframe extraction: One frame is captured every 5 seconds for content analysis.
[0152] Speech-to-text transcription: The ASR engine is called to generate subtitle text and extract technical terms ("coronal high-temperature plasma", "coronal mass ejection").
[0153] Object detection: Identifying observational equipment (such as coronagraph lenses) and astronomical phenomena (such as solar flares) in videos.
[0154] Output supplementary metadata:
[0155] Keyframe timestamps: ["00:00:05", "00:00:10"...]
[0156] List of technical terms: ["Coronal high-temperature plasma", "Coronal mass ejection"]
[0157] Associated resource: ["urn:obs:solar:20231007_CME_001"], which associates with existing observation data.
[0158] 3D Model Analysis Pipeline
[0159] Input: 3D scan file (.obj / .ply), structured metadata (such as scan accuracy, coordinate system standard).
[0160] Processing steps:
[0161] Geometric analysis: Calculate basic physical parameters (surface area, volume, centroid coordinates); detect geometric anomalies (voids, non-manifold structures).
[0162] Multimodal feature extraction: Analyze the surface feature distribution (such as chromaticity and reflectivity gradient) in texture mapping; fuse point cloud density and normal vectors to analyze structural complexity.
[0163] Output supplementary metadata:
[0164] Geometric characteristics: {"Surface area": "2.3m²", "Volume": "0.8m³"}
[0165] Material estimation: {"Metal": "85%", "Non-metal": "12%"} (based on multispectral reflectance analysis)
[0166] Text / Document Parsing Pipeline
[0167] Input: PDF / Word document, structured metadata (author, permission level, document type).
[0168] Processing steps:
[0169] Content enhancement: OCR correction for low-quality scanned documents (such as fuzzy text reconstruction and multilingual mixed recognition); extraction of structured paragraphs (automatic segmentation of titles / body text / figure captions).
[0170] Semantic parsing: cross-domain terminology extraction (technical terms, standard numbers, spatiotemporal descriptions); related entity identification (organization / person / resource identifiers).
[0171] Permission inheritance: Automatically generate derivative resource control policies based on the original permission tags.
[0172] Output supplementary metadata:
[0173] Technical terms: ["spectral analysis", "stratigraphic layering"],
[0174] Related resources: ["specimen:NMNH-2023-001", "experiment:STM-Phase2"],
[0175] Content reuse: {"Edit": "Prohibited", "Quote": "Source must be cited"},
[0176] Based on one example, the final output of the fourth JSON-LD data could contain, for example:
[0177] Metadata: Structured metadata from step S3 + supplementary metadata from step S4 (enhanced information extracted from the original data content through algorithms).
[0178] Processing traceability: Record the pipeline version used, algorithm parameters, and operation timestamps.
[0179] Original data reference: Initial access data is fully preserved for review.
[0180] Type-driven multimodal analysis significantly enhances the knowledge mining capabilities of specialized resources. This domain-knowledge-based algorithm optimization automates signal processing workflows that previously required manual intervention. The dynamic adaptability of the multimodal pipeline effectively addresses the parsing needs of complex resources across multiple venues.
[0181] In step S5, the digital resource acquisition system establishes a correlation between the metadata of different resources based on the fourth JSON-LD data of each access resource, forming the fifth JSON-LD data of each access resource.
[0182] Here, based on the multimodal features of a single resource within the fourth JSON-LD data generated in step S4, the system constructs a metadata association network inside and outside the resource through semantic mapping and association calculation. Each fourth JSON-LD data corresponds to an independent resource entity, such as a collection, an exhibition, or an educational activity, which encapsulates the multimodal metadata and original data references of that resource.
[0183] Metadata association can be based on direct links between explicit rules and structured fields. Its core goal is to quickly establish deterministic relationships between resources through predefined logic. Explicit rules can rely on predefined field mappings and time matching rules in YAML templates. For example, it can rely on precise matching of time fields without semantic analysis.
[0184] Association rules can be designed based on business logic and inherent resource attributes, and the main types include:
[0185] Spatiotemporal association: Resources generated within the same time period or geographical location are automatically associated.
[0186] Topic association: Establish associations between resources that share the same technical terms.
[0187] For example, if a planetarium meteorite analysis report (PDF) mentions meteorite number "METEOR-2024A", the system will perform the following association:
[0188] Resource binding: Associate the report with the multispectral scan file of meteorite number "METEOR-2024A" and the simulation animation of the fall trajectory.
[0189] Thematic aggregation: The report is categorized into the "Meteorite Research" collection and linked to similar meteorite composition databases.
[0190] Metadata association aims to establish semantic and logical connections between different resources, realizing a "networked" organization of science popularization resources. This association is achieved through the following dimensions:
[0191] 1. Topic Relevance
[0192] Extract the "core topic" metadata (such as "photosynthesis" and "black hole formation") of resources from the fourth JSON-LD data. Calculate the topic relevance using a semantic similarity algorithm (such as BERT vector matching). Bind resources with a relevance ≥ 0.7 to each other, record the association relationship using the "relatedTopic" attribute, and label the association weight (e.g., "0.85"). For example, "photosynthesis experiment video" and "plant cell structure diagram" are associated because of their shared topic "photosynthesis".
[0193] 2. Entity Association
[0194] Identify scientific entities in the metadata (such as "oxygen" and "Galileo's telescope"), assign a unique identifier to each entity (such as "Entity ID: O2-001" and "Entity ID: Telescope-003"), and associate resources containing the same entity through the "mentionsEntity" attribute. For example, the text "History of the Discovery of Oxygen" and the 3D model of "Lavoisier's Experimental Apparatus" are associated because of the common entities "oxygen" and "experimental apparatus".
[0195] 3. Origin and Derivative Connections
[0196] If resource B is derived from resource A (e.g., "experimental video clips" originate from "complete experimental recordings"), the original resource is associated through the "derivedFrom" attribute; if multiple resources come from the same acquisition project (e.g., "solar system exploration series resources"), a common source identifier is bound through the "sourceProject" attribute.
[0197] After the association is completed, the fifth JSON-LD data will retain all the metadata of the fourth JSON-LD data and add a "association attribute set" to form a complete structure containing "resource ontology metadata + association relationship metadata", providing a data foundation for resource retrieval and thematic integration for science museums.
[0198] Furthermore, the establishment of a metadata association network for resources enables cross-resource type / cross-modal search and recommendation to significantly surpass traditional unimodal search and recommendation.
[0199] Metadata association networks break down the type barriers and modal boundaries of digital resources in science museums. Their core value lies in transforming scattered resources such as text, images, videos, and 3D models into a semantically interconnected knowledge network, enabling search and recommendation to leap from "based on single keyword matching" to "based on multi-dimensional semantic association." Specific advantages are reflected in the following scenarios:
[0200] 1. Intelligent search across resource types:
[0201] In traditional unimodal search, when a user queries "dinosaur fossils," the search results only return images or text tagged with "dinosaur fossils," and the results are arranged unordered by time / popularity. However, searches based on the association network disclosed herein trigger multi-level associations:
[0202] 1) Directly Associated Layer: Returns "Dinosaur Fossil Specimen Photos" (image resource) + "Fossil Excavation Report" (text resource) + "Fossil Formation Process Animation" (video resource), which are all bound to the same entity "Dinosaur Fossil";
[0203] 2) Theme extension layer: The "relatedTopic" attribute is used to link "Dinosaur Evolution Tree 3D Model" (with a theme relevance of 0.82) and "Cretaceous Geological Profile" (with a theme relevance of 0.76).
[0204] 2. Cross-modal recommendation:
[0205] Traditional recommendation logic often involves "copying similar resources," such as recommending only other space videos after watching "lunar exploration videos." However, recommendations driven by metadata association networks can achieve precise adaptation based on multimodal semantic parsing of user behavior.
[0206] When a user browses a "Quantum Entanglement Diagram" (image resource), the system identifies their implicit need for the "Fundamentals of Quantum Mechanics" topic through the "relatedTopic" attribute and makes recommendations accordingly.
[0207] 1) Text resource: "A Popular Explanation of Quantum Entanglement" (with supplementary theoretical background);
[0208] 2) Video resources: "Quantum entanglement experiment simulation video" (matching dynamic processes not shown in the schematic diagram via the "complements" attribute).
[0209] 3. Core competencies that transcend tradition:
[0210] The underlying logic of metadata association networks is the semantic standardization of metadata (achieved through JSON-LD data) and the quantitative calculation of association weights (such as topic relevance and entity co-occurrence frequency), which gives it capabilities that are difficult for traditional systems to achieve:
[0211] 1) Fuzzy query processing: When a user enters "Why do stars shine?" (natural language), the system extracts the query terms "star" and "nuclear fusion" through text parsing. It not only returns the corresponding resources, but also related resources such as "stellar spectrum" (image), "nuclear fusion simulation video" (video), and "stellar evolution paper" (text), without requiring the user to precisely match the tags;
[0212] 2) Knowledge transfer: When a user queries "Earth's rotation", the system not only returns "Earth's rotation animation" (video), but also returns supplementary resources identified by the "associated attribute set", such as "day and night alternation interactive game", to achieve "knowledge transfer" across resources.
[0213] This search and recommendation based on metadata-related networks essentially transforms individual science popularization resources into a "knowledge network," which not only lowers the threshold for the public to access cross-domain information but also provides data support for the planning of special exhibitions and the design of educational activities in science museums. For example, it can quickly integrate cross-modal resource packages such as papers, experimental videos, and data visualization charts on the theme of "carbon."
[0214] It should be noted that the embodiments of this disclosure can be implemented in software and / or a combination of software and hardware, for example, using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In one embodiment, the software program of the embodiments of this disclosure can be executed by a processor to implement the steps or functions described above. Similarly, the software program (including associated data structures) of the embodiments of this disclosure can be stored in a computer-readable recording medium, such as RAM memory, a magnetic or optical drive, a floppy disk, or similar devices. Furthermore, some steps or functions of the embodiments of this disclosure can be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.
[0215] Furthermore, at least a portion of the embodiments of this disclosure can be applied as computer program products, such as computer program instructions, which, when executed by a computing device, can invoke or provide methods and / or technical solutions according to the embodiments of this disclosure through the operation of the computing device. The program instructions that invoke / provide the methods of the embodiments of this disclosure may be stored in a fixed or removable recording medium, and / or transmitted via a data stream in a broadcast or other signal carrying medium, and / or stored in the working memory of a computing device operating according to the program instructions.
[0216] It will be apparent to those skilled in the art that the embodiments of this disclosure are not limited to the details of the exemplary embodiments described above, and that the embodiments of this disclosure can be implemented in other specific forms without departing from the spirit or essential characteristics of the embodiments of this disclosure. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the embodiments of this disclosure is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be encompassed within the embodiments of this disclosure. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is apparent that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in the system claims may also be implemented by a single unit or device in software or hardware. The terms "first," "second," etc., are used to denote names and do not indicate any particular order.
Claims
1. A method for acquiring digital resources for science museums, wherein, The method includes the following steps: Access multi-source heterogeneous data from science museums, including collections, exhibitions, or educational activities, and generate first JSON-LD data for each accessed data, including standardized basic metadata and original feature metadata; Based on the first JSON-LD data, explicit technical features are extracted from the original feature metadata to generate a first type label, and semantic information in the standardized basic metadata is parsed to generate a second type label. The first type label and the second type label are then merged to obtain the final type label, forming the second JSON-LD data. Based on the second JSON-LD data, a predefined rule template is loaded according to the type tag to extract structured metadata, including type core metadata and type-specific metadata, while retaining the reference to the original access data to form the third JSON-LD data; Based on the third JSON-LD data, the corresponding pipeline branch is determined from the multimodal parsing pipeline, which includes text / document parsing pipeline, video / audio parsing pipeline and 3D model parsing pipeline, according to the type label. The original access data is parsed using the pipeline branch to generate supplementary metadata. The structured metadata provides domain prior knowledge for the pipeline branch when parsing the original access data, optimizes the algorithm parameters, and forms the fourth JSON-LD data. Based on the fourth JSON-LD data of each access data, a relationship is established between the metadata of different access data to form the fifth JSON-LD data of each access data. The standardized basic metadata consists of common fields for all resource types, containing only business semantic fields shared by all resources and independent of specific technical implementations; the original feature metadata records the technical characteristics and access details of the resources and is stored according to data source type. The fusion of the first type of tag and the second type of tag includes any of the following: - Determine the final type label based on the weighted confidence scores of the first type label and the second type label; - When the confidence level of the second type label exceeds a predetermined threshold, the second type label is used as the final type label; - Use the first type of tag as the parent class and the second type of tag as the child class to form a hierarchical type tag.
2. The method according to claim 1, wherein, The first type of label includes the basic type of rule triggering, and the second type of label includes the subclass of model recognition.
3. The method according to claim 1, wherein, The association is based on field matching according to predetermined association rules.
4. The method according to claim 1 or 3, wherein, The associations of the metadata include subject associations, entity associations, or source-derived associations.
5. The method according to claim 1, wherein, The method also includes the following steps: Based on the establishment of the metadata association network for each access data, recommendations / searches across resource types / modalities are provided.
6. A digital resource acquisition system for science museums, wherein, The system includes a memory and a processor. The memory stores computer program instructions, which, when executed by the processor, configure the system to perform the following operations: Access multi-source heterogeneous data from science museums, including collections, exhibitions, or educational activities, and generate first JSON-LD data for each accessed data, including standardized basic metadata and original feature metadata; Based on the first JSON-LD data, explicit technical features are extracted from the original feature metadata to generate a first type label, and semantic information in the standardized basic metadata is parsed to generate a second type label. The first type label and the second type label are then merged to obtain the final type label, forming the second JSON-LD data. Based on the second JSON-LD data, a predefined rule template is loaded according to the type tag to extract structured metadata, including type core metadata and type-specific metadata, while retaining the reference to the original access data to form the third JSON-LD data; Based on the third JSON-LD data, the corresponding pipeline branch is determined from the multimodal parsing pipeline, which includes text / document parsing pipeline, video / audio parsing pipeline and 3D model parsing pipeline, according to the type label. The original access data is parsed using the pipeline branch to generate supplementary metadata. The structured metadata provides domain prior knowledge for the pipeline branch when parsing the original access data, optimizes the algorithm parameters, and forms the fourth JSON-LD data. Based on the fourth JSON-LD data of each access data, a relationship is established between the metadata of different access data to form the fifth JSON-LD data of each access data. The standardized basic metadata consists of common fields for all resource types, containing only business semantic fields shared by all resources and independent of specific technical implementations; the original feature metadata records the technical characteristics and access details of the resources and is stored according to data source type. The fusion of the first type of tag and the second type of tag includes any of the following: - Determine the final type label based on the weighted confidence scores of the first type label and the second type label; - When the confidence level of the second type label exceeds a predetermined threshold, the second type label is used as the final type label; - Use the first type of tag as the parent class and the second type of tag as the child class to form a hierarchical type tag.
7. A computer program product comprising computer program instructions, wherein, When the computer program instructions are executed by a computer device, the computer device is configured to perform the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Label description method and device for multisource isomerism scientific and technical information recourses
CN107357933A