A method for constructing a cultural resource big data reference model and application thereof
By combining thematic and tag domains and employing big data and multimodal fusion technologies, a big data reference model for cultural resources is constructed. This addresses the shortcomings of existing models in personalized services and data utilization, enabling the model to be segmented and diversified, and improving the efficiency of cultural resource sharing and recommendation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NORTHWEST UNIV
- Filing Date
- 2022-08-05
- Publication Date
- 2026-04-14
AI Technical Summary
Existing methods for constructing big data reference models for cultural resources fail to fully integrate subject and tag domains, making it difficult to achieve personalized services and effectively utilize multiple data types, resulting in a lack of diversity and flexibility in model applications.
By adopting a thematic domain combined with tag domain approach, and combining big data and multimodal thinking, a big data reference model for cultural resources is constructed by subdividing horizontal and vertical data models. Then, big data architecture and multimodal fusion technology are used to realize the diversified expression and application of data.
It has achieved a detailed breakdown of the model's main framework, provided personalized services, optimized business processes, improved the diversity of online and offline applications and data utilization efficiency, and met the needs of cultural resource sharing and recommendation for different purposes.
Smart Images

Figure CN115309838B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of cultural big data technology, specifically relating to a method for constructing and applying a cultural resource big data reference model. Background Technology
[0002] Since the beginning of the new century, new-generation information technologies such as the Internet of Things, cloud computing, big data, and artificial intelligence have been used to comprehensively transform public cultural services and the cultural industry across the entire chain, continuously promoting the networking and intelligentization of public cultural construction, while placing greater emphasis on providing personalized and precise services to users. Through thorough research into the current state of digital development of public culture in China, this paper summarizes and concludes the complete process of public cultural digital resources from generation to application, and establishes a big data reference model for cultural digital resources. This model mainly includes the definition, categories, lifecycle, involved roles and management, resource description requirements, and reference model framework of public cultural digital resources. Current research on the classification of cultural big data can be based on two perspectives: classification based on data content themes and classification based on data dimensions.
[0003] With the continuous development of information technology, the application of cultural resources has rapidly expanded to various fields, making the construction of big data models for cultural resources and the strengthening of digitalization an inevitable trend. Public cultural cloud refers to fully utilizing cloud computing technology to integrate various digital cultural resources and, based on this, using network technology to build a service platform to achieve universal sharing of cultural resources. Public cultural cloud can effectively improve the user experience of cultural centers, diversify service channels, and thus enhance the practicality and utilization of cultural centers; it represents an effective integration and unification of the internet and big data.
[0004] With my country's increasing emphasis on cultural soft power and the continuous development of information technology, strengthening the construction of big data models for cultural resources has become a top priority in my country's cultural dissemination efforts. The rational application of these models is a key measure in the digital construction of cultural centers at all levels, an important way to enrich the cultural life of grassroots communities, and an effective unification of traditional culture and modern technology. In the digital construction of culture, the public is not only a beneficiary of culture but also a disseminator of it, greatly promoting the sharing of cultural resources. Summary of the Invention
[0005] In order to overcome the shortcomings of the prior art, the present invention aims to provide a method for constructing and applying a big data reference model for cultural resources. Unlike traditional model construction methods, the construction method proposed in this invention combines big data and multimodal thinking on the basis of combining thematic domains and tag domains.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0007] A method for constructing a big data reference model for cultural resources, characterized by the following steps:
[0008] Step 1: Define a reference model for big data on cultural resources, establish the model domain and model foundation, and further subdivide the domain's horizontal data model and vertical data model;
[0009] Step Two: Constructing the Main Body of a Big Data Reference Model for Cultural Resources by Combining Thematic Domains with Tag Domains:
[0010] The tag domain reference model was further subdivided by combining the topic domain model. Taking into account the index storage involved in the platform file system in the interactive system, the source tag domain reference model was constructed. It is a collection of concepts related to the cultural resource tag domain, reflecting the relationship between the various concepts and the rules for managing interactions.
[0011] Step 3: Combine the model built by integrating the main domain with the tag domain with big data:
[0012] Big data architecture can be used to represent stacked or chained systems composed of multiple big data systems, where a data consumer in one system can act as a data provider for a later system. Combining big data architecture with a cultural resource big data reference model helps people understand how big data systems can supplement existing systems and differ from traditional data application systems such as analysis, business intelligence, and databases. It is precisely because of the excellent characteristics of big data that big data thinking runs through the entire model building process.
[0013] Step 4: Further integrate multimodal thinking after step 3:
[0014] Transforming different forms of data into a unified form of data expression through multimodal alignment, multimodal fusion, and multimodal joint representation, thereby realizing its semantic meaning and fully realizing the utility of big data;
[0015] Step 5, Application of the reference model:
[0016] The cultural resource big data reference model is applied to online and offline systems. During the construction process, a user recommendation-related application reference model and an application platform integration reference model are also generated, thus applying the cultural resource big data reference model in practice.
[0017] The aforementioned model foundation includes a general horizontal cultural resource data model and terminology and data dictionary for the cultural resource field, and on this basis, data models for three types of cultural resource fields are established: library field, museum field, and cultural center field.
[0018] The construction of the subject domain of the data model is specifically broken down into the following steps:
[0019] The data model for the library domain includes structured horizontal models and unstructured horizontal models for the book domain. The structured horizontal models for the book domain are further divided into four types of vertical models: layout, cover, table of contents, and content. The unstructured horizontal models for the book domain are further divided into three types of vertical models: tag, comment, and summary.
[0020] The museum domain data model includes a horizontal model for the museum domain, which is then divided into four types of vertical models: vertical models for the history domain, vertical models for the art domain, vertical models for the natural sciences domain, and vertical models for comprehensive domains.
[0021] The data model for the cultural center sector includes a structured horizontal model for the cultural sector, which is then divided into four vertical models: a comprehensive vertical model, a non-material cultural heritage vertical model, an ethnic cultural vertical model, and a digital vertical model.
[0022] The thematic domains can be further subdivided vertically into sections such as Cultural Cloud, Artists, and Cultural Centers, Museums, and Libraries. Horizontal thematic domains can be divided in correspondence with the vertical thematic domains. The Artists domain corresponds to the User domain; the integrated Cultural Centers domain corresponds to the domain of artists, cultural relic guides, and librarians, which is related to the User domain and can be further subdivided into thematic tag domains; the Cultural Cloud domain corresponds to various horizontal activity subdivisions.
[0023] The topic domain model is further subdivided into label domains, and the specific subdivision steps are as follows:
[0024] The tag domain consists of the tag operation domain and the tag type domain;
[0025] The tag operation domain includes four types of domain-based horizontal data models: tag filtering horizontal data model, tag deduplication horizontal data model, tag merging horizontal data model, and weight sorting horizontal data model;
[0026] The tag type domain mainly includes two categories: user tag domain and resource tag domain. These can be further divided into four horizontal data models: person tag horizontal data model, category tag horizontal data model, object / scene tag horizontal data model, and text tag horizontal data model. The person tag horizontal data model is further divided into three vertical data models: face vertical data model, expression vertical data model, and character vertical data model. The category tag horizontal data model is divided into two vertical data models: action / event vertical data model and organization vertical data model. The object / scene tag horizontal data model is divided into two vertical data models: scene vertical data model and object vertical data model. The text tag horizontal data model is divided into two vertical data models: region vertical data model and identifier vertical data model.
[0027] The model described incorporates big data thinking, specifically as follows:
[0028] The Cultural Resources Big Data Reference Model utilizes big data thinking. In the process of cultural resource data processing, the first step is data collection, which is then fed into the overall framework of the cultural cloud. Cultural resources are categorized and data is graded, and then horizontal domains such as libraries, museums, and cultural centers are divided. These horizontal domains form vertical thematic domains, which are further modeled to build a domain knowledge model. Finally, data destruction is performed. The entire process involves data storage, processing, transmission, and exchange. Big data is also linked to cultural resources. The model starts at the data collection layer, which is divided into movable and immovable cultural heritage data. This data is then fed into the metadata description layer, where visualized resources and other resources are standardized using metadata. The standardized metadata ontology is then split into two different ontology layers at the ontology construction layer. These ontology layers are then fed into the association data layer to associate various resources. Finally, at the intelligent application layer, applications are promoted through associated data generation, publication, browsing, querying, and intelligent recommendation and matching. This generates sub-models of the Cultural Resources Big Data Reference Model, including a user recommendation-related application reference model and an application platform integration application model.
[0029] The aforementioned big data thinking integrates multimodal concepts, specifically as follows:
[0030] The cultural resource big data reference model is implemented using the concept of a multimodal model. In this model, data is categorized into formatted and unformatted types. Formatted data is represented by tables, while unformatted data includes images, tags, files, and audio. Data is transformed into different embedding vectors for multimodal alignment. Different multimodal processing methods are used: single-stream multimodal and dual-stream multimodal. Single-stream multimodal processing takes text, images, or audio information as separate inputs and concatenates them for output. Dual-stream multimodal processing takes text and images as simultaneous inputs and outputs them together. Single-stream and dual-stream multimodal processing are collectively referred to as multimodal joint representation. The processed data is aggregated through fusion multimodal processing. A score table generated by multiplying multimodal features by semantic vectors of classification terms is compared with a classification data table. Each feature vector is then classified into the category with the highest score, thus achieving classification.
[0031] The tagging model also incorporates multimodal thinking, designing tags that encompass various concepts within the domain and reflecting the relationships between tags (the rules and forms of tag division). By analyzing the applicability of tags, it clarifies the scope and degree of data tag reuse from a data perspective, thus classifying tags into: text tags, image tags, audio tags, and video tags. Domain-level horizontal entities are tag data that are relatively universal within the domain; domain-level vertical entities refer to those applicable only to a specific domain, not referenced by domain-level horizontal tag data, and their relationships with other domain-level vertical tag data are often definite. Therefore, in unimodal tagging, domain-level horizontal entities can be divided into text tags, image tags, audio tags, and video tags, while independent vertical entity tags within their respective domains can be classified separately. For example, text tags include regions and keywords; image tags include faces and objects; audio tags include songs, dances, movies, and dances; and video tags include long videos and short videos. In multimodal tagging, domain-level horizontal entities can be divided into multimodal alignment, multimodal fusion, and multimodal joint representation.
[0032] The application of the cultural resource big data reference model specifically refers to its application in online and offline systems. The framework for public cultural digital resources should effectively combine online and offline approaches. The application can be further subdivided into the application of library data resources, museum data resources, public cultural cloud data resources, and cultural center data resources.
[0033] The beneficial effects of this invention are:
[0034] Compared with existing methods for constructing cultural resource big data reference models, the cultural resource big data reference model of this invention combines thematic domains with tag domains, which can better achieve the subdivision of the model's main structure, comprehensively consider the model's positioning, and provide personalized services for different needs. At the same time, the reference model also incorporates big data and multimodal considerations, enabling the model to utilize different types of data and better leverage complex data. This allows the model to better understand and optimize business processes, establish user recommendation-related application reference models and application platform integration reference models, making its online and offline applications more diverse. Attached Figure Description
[0035] Figure 1 A diagram for constructing a big data reference model for cultural resources;
[0036] Figure 2 Recommend relevant application reference models to users;
[0037] Figure 3 A platform integration reference model for applications;
[0038] Figure 4 As a reference model for the label domain;
[0039] Figure 5 Interactive diagrams for horizontal and vertical subdivision of thematic areas. Detailed Implementation
[0040] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0041] like Figures 1 to 5 As shown, a method for constructing a cultural big data resource reference model based on cultural big data includes the following steps:
[0042] Step one: Define the cultural resource big data reference model and establish the model domain. The cultural resource domain data reference model conforms to the "divide and conquer" principle. The system architecture provides a framework and basic guidance for data modeling, while also facilitating different degrees of reuse and overall model upgrade planning. Its model foundation includes: a general horizontal cultural resource data model and a terminology and data dictionary for the cultural resource domain. Based on this, three types of data models for cultural resource domains are established: library domain, museum domain, and cultural center domain. These are further subdivided into horizontal and vertical data models for each domain.
[0043] Step Two: Constructing a Cultural Resource Big Data Reference Model by Combining Theme Domains and Tag Domains. The main body of the model is the tag domain reference model, which is further subdivided based on the theme domain model. Taking into account the index storage involved in the platform file system in the interactive system, a source tag domain reference model is constructed. This model is a collection of related concepts in the cultural resource tag domain, reflecting the relationships between various concepts and the rules for managing interactions.
[0044] Step three involves integrating the model constructed from the subject domain and the labeled domain with big data: Big data architecture can be used to represent stacked or chained systems composed of multiple big data systems, where a data consumer in one system can act as a data provider for a subsequent system. Combining big data architecture with a cultural resource big data reference model helps people understand how big data systems complement existing systems and differ from traditional data application systems such as analytics, business intelligence, and databases. It is precisely because of the favorable characteristics of big data that big data thinking permeates the entire model construction process.
[0045] Step four, following step three, further integrates multimodal thinking: transforms data of different forms into a unified form of data expression through multimodal alignment, multimodal fusion, and multimodal joint representation, so that it is semantically grounded and the utility of big data is fully realized.
[0046] Step 5, Application of the Reference Model: The cultural resource big data reference model is applied to online and offline systems. During the construction process, a user recommendation-related application reference model and an application platform integration reference model are also generated, thus applying the cultural resource big data reference model in practice.
[0047] It can be further divided into the following stages:
[0048] Phase One: Defining a Big Data Reference Model for Cultural Resources. First, the model domain is established. The data reference model for the cultural resource domain conforms to the "divide and conquer" principle. Its architecture provides a framework and basic guidance for data modeling, while also facilitating different degrees of reuse and overall model upgrade planning. Its model foundation includes: a general horizontal cultural resource data model and a terminology and data dictionary for the cultural resource domain. Based on this, three types of data models are established for the cultural resource domain: library domain, museum domain, and cultural center domain. Further subdivisions are made into horizontal and vertical data models for each domain.
[0049] The library domain data model includes both structured and unstructured horizontal models for the book domain. The structured horizontal models are further divided into four vertical domain models: layout, cover, table of contents, and content. The unstructured horizontal models are divided into three vertical domain models: tag, review, and summary.
[0050] The museum domain data model includes a horizontal model of the museum domain, and based on this model, it is divided into four types of vertical models: vertical models of the history domain, vertical models of the art domain, vertical models of the natural science domain, and vertical models of the comprehensive domain.
[0051] The data model for the cultural center sector includes a structured horizontal model for the cultural sector, which is then divided into four vertical models: a comprehensive vertical model, a non-material cultural sector vertical model, an ethnic sector vertical model, and a digital sector vertical model.
[0052] Further subdividing the thematic areas, the vertical thematic areas can be divided into sections encompassing cultural cloud, artists and writers, and cultural centers, museums, and libraries. The horizontal thematic areas can correspond to these vertical sections. The artist and writer section corresponds to the user section; the integrated section encompassing cultural centers, artists and writers, cultural relic guides, and librarians, and is linked to the user section, further subdivided into thematic tag sections; the cultural cloud section corresponds to various horizontal activity sub-sections, such as: reading good books, gathering information, gathering literary talent, and booking venues. These various sections interact with each other, forming an interactive system.
[0053] Phase Two: Constructing the Main Body of a Cultural Resource Big Data Reference Model by Combining Theme Domains and Tag Domains. The tag domain reference model is further subdivided based on the theme domain model. Considering the index storage involved in the platform file system within the interactive system, a source tag domain reference model is constructed. This model is a collection of concepts related to the cultural resource tag domain, reflecting the relationships between these concepts and the rules for managing interactions. It consists of a tag operation domain and a tag type domain.
[0054] The tag operation domain includes four types of domain-based horizontal data models: tag filtering horizontal data model, tag deduplication horizontal data model, tag merging horizontal data model, and weight sorting horizontal data model.
[0055] The tag type domain mainly comprises two categories: user tag domain and resource tag domain, which can be further divided into four horizontal data models: person tag horizontal data model, category tag horizontal data model, object / scene tag horizontal data model, and text tag horizontal data model. Within the person tag horizontal data model, there are three vertical data models: face vertical data model, expression vertical data model, and character vertical data model; within the category tag horizontal data model, there are two vertical data models: action / event vertical data model and organization vertical data model; within the object / scene tag horizontal data model, there are two vertical data models: scene vertical data model and object vertical data model; and within the text tag horizontal data model, there are two vertical data models: geographical vertical data model and identifier vertical data model.
[0056] Phase Three: The Cultural Resource Big Data Reference Model Utilizes Big Data Thinking. In the cultural resource data processing, data collection begins, and the data is sent to the overall framework of the cultural cloud for classification and grading. Then, the data is divided into horizontal domains such as libraries, museums, and cultural centers. These horizontal domains form vertical thematic domains, which are further modeled to build a domain knowledge model. Finally, the data is destroyed. The entire process involves data storage, processing, transmission, and exchange. Big data is linked to cultural resources; the model starts at the data collection layer, dividing it into movable and immovable cultural heritage data. The data is then sent to the metadata description layer, where visualized resources and other resources are standardized using metadata. The standardized metadata ontology is then split into two different ontologies at the ontology construction layer. These ontologies are sent to the association data layer to associate various resources. Finally, at the intelligent application layer, applications are promoted through associated data generation, publication, browsing, querying, and intelligent recommendation and matching. This generates sub-models of the cultural resource big data reference model, including a user recommendation application reference model and an application platform integration application model.
[0057] The user recommendation-related application reference model, referencing the Big Data Standard (BDRA), unfolds from two value chains representing different dimensions of big data organization: the information value chain (horizontal axis) and the information technology value chain (vertical axis). The main role of the information value chain is to demonstrate the information flow value realized by big data in the process of information processing from data to knowledge, serving as a data science method. Its core value is realized through activities such as data collection, processing, user interest prediction, relevant resource prediction, and personalized recommendations. The information technology value chain, as an emerging data application paradigm, reflects the value brought by big data in response to new demands generated by information technology. Its core value is realized through providing networks for storing and running big data applications, prediction and recommendation models for the user side, resource side, and other information technology services. Public cultural cloud resource data is located at the intersection of these two value chains, employing big data analytics to provide specific value to big data stakeholders across both value chains.
[0058] BDRA also provides a component hierarchy classification system to describe the logical components in BDRA and define their classification. Logical components in BDRA are divided into three levels, from highest to lowest: roles, activities, and components. The top-level logical components, roles, represent the five roles existing in the big data system: System Coordinator (ETL), Data Provider (User), Big Data Application Provider (Resource Data), Big Data Framework Provider (Model Management), and Data Consumer (System Data Security). The second-level logical components are the activities performed by each role. The third-level logical components are the functional components needed to perform each activity, such as intelligent analysis and management, which provide services and functions for the five roles in the big data system.
[0059] Similarly, the platform integration reference model for applications unfolds from two value chains representing different dimensions of big data organization: the information value chain (horizontal axis) and the information technology value chain (vertical axis). The core value of the information value chain is realized through activities such as Flume, acquisition, collection, solidification, and evaluation. The core value of the information technology value chain is realized by providing networks for storing and running big data, prediction and recommendation models for users and resources, and other information technology services. Zookeeper, Yarn, and Elastic Search are located at the intersection of these two value chains, employing big data analytics and providing specific value to big data stakeholders across both chains.
[0060] The component hierarchy classification system includes five roles: system coordinator (administrator), data provider (user), big data application provider (Zookeeper, Yarn, and Elastic Search), big data framework provider (hyperconverged infrastructure), and data consumer (application system). The second-level logical components are the activities performed by each role. The third-level logical components are the functional components required to perform each activity, such as subsystem web interfaces and management, which provide services and functions for the five roles of the big data system.
[0061] Big data architecture can be used to represent stacked or chained systems composed of multiple big data systems, where a data consumer in one system can act as a data provider for a subsequent system. Combining big data architecture with a cultural resource big data reference model helps people understand how big data systems complement existing systems and differ from traditional data application systems such as analytics, business intelligence, and databases. It is precisely because of the favorable characteristics of big data that big data thinking permeates the entire model building process.
[0062] Phase Four: The Cultural Resource Big Data Reference Model implements big data through a multimodal model. In this model, data is categorized into formatted and unformatted types. Formatted data is represented by tables, while unformatted data includes images, tags, files, and audio. Data is transformed into different embedding vectors for multimodal alignment. Different multimodal processing methods are used: single-stream multimodal and dual-stream multimodal. Single-stream multimodal processing takes text, images, or audio as separate inputs and concatenates them for output. Dual-stream multimodal processing takes both text and images as inputs and outputs them together. Single-stream and dual-stream multimodal processing are collectively referred to as multimodal joint representation. The processed data is aggregated through fusion multimodal processing. A score table generated by multiplying multimodal features by semantic vectors of classification terms is compared with a classification data table. Each feature vector is then classified into the category with the highest score, thus achieving classification.
[0063] The tagging model also incorporates multimodal thinking, designing tags that encompass various concepts within the domain and reflecting the relationships between tags (the rules and forms of tag division). By analyzing the applicability of tags, the scope and degree of data tag reuse are clarified from a data perspective, thus classifying tags into: text tags, image tags, audio tags, and video tags. Domain-level horizontal entities are tag data that are relatively universal within the domain; domain-level vertical entities refer to those applicable only to a specific domain, not referenced by domain-level horizontal tag data, and their relationships with other domain-level vertical tag data are often definite. Therefore, in unimodal tagging, domain-level horizontal entities can be divided into text tags, image tags, audio tags, and video tags, while independent vertical entity tags within their respective domains can be classified separately. For example, text tags include regions and keywords; image tags include faces and objects; audio tags include songs, dances, movies, and dances; and video tags include long videos and short videos. In multimodal tagging, domain-level horizontal entities can be divided into multimodal alignment, multimodal fusion, and multimodal joint representation.
[0064] Phase Five: Application of the Cultural Resource Big Data Reference Model to Online and Offline Systems. Online and offline cultural products are essentially inseparable; their differences lie primarily in their medium and form of expression. Offline cultural resources offer the advantage of being tangible and interactive, providing a strong sense of immersion and diversity. Online cultural resources, on the other hand, offer convenient access, long-term preservation, repeatable consumption, and high capacity. During the pandemic, with the closure of offline cultural venues and the cancellation of various public cultural activities, online services relying on internet technology became a new choice for people's cultural lives. Therefore, the framework for public cultural digital resources should grasp the concept of effectively combining online and offline resources, recognizing that they are no longer separate but effectively integrated and unified, even fused into one. The application of this reference model can be further subdivided into the application of library data resources, museum data resources, public cultural cloud data resources, and cultural center data resources.
[0065] Example:
[0066] First, a model building environment based on topic tags is established, including the division of the model's main topic modules and the sub-categories of tags within them. Specifically, the domains of libraries, museums, and cultural centers are defined first. Then, diverse and complex data are aggregated using big data, and multimodal thinking is integrated to transform different forms of data into a unified data representation, thus grounding it in semantics. Finally, based on practical applications, the tag domain is further subdivided from broad categories of user tags and resource tags into short tags, image description tags, title description tags, automatically generated layout tags, and multimodal tags. Tag filtering, tag deduplication, tag merging, and weighted ranking are used to apply the final model to specific online and offline systems, specifically as a reference model for user recommendation applications and a reference model for platform integration.
[0067] This method leverages the advantages of big data and multimodal approaches, along with the innovative approach of combining topic and tag domains. Addressing the shortcomings of previous model building methods, this invention, by combining topics and tags, can comprehensively consider different needs, greatly satisfying the public's cultural experience requirements. Furthermore, by combining big data and multimodal approaches, it can better utilize complex data to accurately locate requirements, maximizing the value of current cultural data.
Claims
1. A method for constructing a big data reference model for cultural resources, characterized in that, Includes the following steps: Step 1: Define a reference model for big data on cultural resources, establish the model domain and model foundation, and further subdivide the domain's horizontal data model and vertical data model; Step Two: Constructing the Main Body of a Big Data Reference Model for Cultural Resources by Combining Thematic Domains with Tag Domains: The tag domain reference model was further subdivided by combining the topic domain model. Considering the index storage involved in the platform file system in the interactive system, the source tag domain reference model was constructed. It is a collection of concepts related to the cultural resource tag domain, reflecting the relationship between the concepts and the rules for managing the interaction. Step 3: Combine the model built with the topic domain and tag domain with big data: Big data architecture is used to represent a stacked or chained system composed of multiple big data systems, where the data consumer of one system acts as the data provider of the next system. The tag domain reference model is further subdivided by combining the topic domain model. Considering the index storage involved in the platform file system in the interactive system, the source tag domain reference model is constructed. It is a collection of concepts related to the cultural resource tag domain, reflecting the relationship between the concepts and the rules for managing the interaction. Step 4: After step 3, transform different forms of data into a unified form of data expression through multimodal alignment, multimodal fusion, and multimodal joint representation, so that it is semantically grounded, and the utility of big data can be fully realized to achieve further integration of multimodal thinking; Step 5, Application of the reference model: The cultural resource big data reference model is applied to online and offline systems. During the construction process, a user recommendation-related application reference model and an application platform integration reference model are also generated, thus applying the cultural resource big data reference model in practice. The model domain described is a data reference model for the cultural resources domain, which conforms to the idea of "divide and conquer". The system architecture provides a framework and basic guidance for data modeling, while also facilitating different degrees of reuse and the upgrading plan of the entire model. The aforementioned model foundation includes a general horizontal cultural resource data model and terminology and data dictionary for the cultural resource field, and on this basis, data models for three types of cultural resource fields are established: library field, museum field, and cultural center field. The aforementioned domain-specific horizontal data models include a tag filtering horizontal data model, a tag deduplication horizontal data model, a tag merging horizontal data model, and a weighted sorting horizontal data model; The aforementioned domain vertical data models include: face vertical data model, expression vertical data model, and character vertical data model.
2. The method for constructing a cultural resource big data reference model according to claim 1, characterized in that, The construction of the subject domain of the data model is specifically broken down into the following steps: The data model for the library domain includes structured horizontal models and unstructured horizontal models for the book domain. The structured horizontal models for the book domain are further divided into four types of vertical models: layout, cover, table of contents, and content. The unstructured horizontal models for the book domain are further divided into three types of vertical models: tag, comment, and summary. The museum domain data model includes a horizontal model for the museum domain, which is then divided into four types of vertical models: vertical models for the history domain, vertical models for the art domain, vertical models for the natural sciences domain, and vertical models for comprehensive domains. The data model for the cultural center sector includes a structured horizontal model for the cultural sector, which is then divided into four vertical models: a comprehensive vertical model, a non-material cultural heritage vertical model, an ethnic cultural vertical model, and a digital vertical model. The thematic areas are further subdivided. Vertically, the thematic areas are divided into sections such as Cultural Cloud, Literary and Artistic Creators, and Cultural Centers, Museums, and Libraries. Horizontally, the thematic areas are divided in accordance with the vertical thematic areas. The Literary and Artistic Creators section corresponds to the User section. The section that integrates Literary and Artistic Propaganda Workers, Cultural Relics Guides, and Librarians is the Cultural Center section, which is related to the User section and further subdivided into thematic tag areas. The Cultural Cloud section corresponds to various horizontal activity subdivisions.
3. The method for constructing a cultural resource big data reference model according to claim 1, characterized in that, The topic domain model is further subdivided into label domains, and the specific subdivision steps are as follows: The tag domain consists of the tag operation domain and the tag type domain; The tag operation domain includes four types of domain-based horizontal data models: tag filtering horizontal data model, tag deduplication horizontal data model, tag merging horizontal data model, and weight sorting horizontal data model; The tag type domain comprises two main categories: user tag domain and resource tag domain. These are further divided into four horizontal data models: person tag horizontal data model, category tag horizontal data model, object / scene tag horizontal data model, and text tag horizontal data model. The person tag horizontal data model is further subdivided into three vertical data models: face vertical data model, expression vertical data model, and character vertical data model. The category tag horizontal data model is further subdivided into two vertical data models: action / event vertical data model and organization vertical data model. Based on the horizontal data model of object scene labels, two types of domain vertical data models are divided: scene vertical data model and object vertical data model; based on the horizontal data model of text labels, two types of vertical data models are divided: region vertical data model and identifier vertical data model.
4. The method for constructing a cultural resource big data reference model according to claim 1, characterized in that, The model described incorporates big data thinking, specifically as follows: The Cultural Resources Big Data Reference Model utilizes big data thinking. In the process of cultural resource data processing, the first step is data collection, which is then fed into the overall framework of the cultural cloud. Cultural resources are categorized and data is graded, and then horizontal domains such as libraries, museums, and cultural centers are divided. Further horizontal and vertical entity relationship modeling is then conducted to build a domain knowledge model. Finally, data destruction is performed. The entire process involves data storage, processing, transmission, and exchange. Big data is also linked to cultural resources. The model starts at the data collection layer, which is divided into movable and immovable cultural heritage data. This data is then fed into the metadata description layer, where visualized resources and other resources are standardized using metadata. The standardized metadata ontology is then split into two different ontology layers at the ontology construction layer. These ontology layers are then fed into the association data layer to associate various resources. Finally, at the intelligent application layer, applications are promoted through associated data generation, publication, browsing, querying, and intelligent recommendation and matching. This generates sub-models of the Cultural Resources Big Data Reference Model, recommends relevant application reference models, and integrates application platform models.
5. The method for constructing a cultural resource big data reference model according to claim 1, characterized in that, Big data thinking integrates multimodal thinking, specifically: The cultural resource big data reference model is implemented by applying the concept of a multimodal model to big data. In the multimodal model, data is divided into formatted and unformatted types. Formatted data is represented by tables, while unformatted data is represented by images, tags, files, and audio. Data is transformed into different embedding vectors for multimodal alignment, and different multimodal processing methods are divided into single-stream multimodal and dual-stream multimodal processing. Single-stream multimodal learning takes text, image, or audio information as inputs and feeds them into the model separately, then concatenates them to produce the output. Dual-stream multimodal learning takes text and image information as inputs simultaneously and feeds them into the model to produce the output. Single-stream multimodal learning and dual-stream multimodal learning are collectively referred to as multimodal joint representation. The processed data information is aggregated through fusion multimodal learning. The score table generated by multiplying the multimodal features by the semantic vector of the classification word is compared with the classification data table, and each feature vector is classified into the category with the highest score, thereby achieving the classification effect. The tagging model also incorporates multimodal thinking, designing tags that include various concepts within the domain and reflecting the relationships between tags, i.e., the rules and forms of tag division. By analyzing the applicability of tags, the scope and degree of data tag reuse are clarified from a data perspective, thus dividing tags into: text tags, image tags, audio tags, and video tags. Domain-level horizontal entities are tag data that are relatively universal within the domain; domain-level vertical entities refer to those that are only applicable to a specific domain, are not referenced by domain-level horizontal tag data, and have a definite relationship with vertical tag data in other domains. Therefore, in unimodal tagging, domain-level horizontal entities are divided into text tags, image tags, audio tags, and video tags, while independent vertical entity tags in their respective domains are classified separately, i.e., region and keywords in the text tag domain; and faces and objects in the image tag domain. The audio tag category includes songs and dances, movies and dances; the video tag category includes long videos and short videos. Multimodal domain horizontal entities can be categorized into multimodal alignment, multimodal fusion, and multimodal joint representation.
6. The method for constructing a big data reference model for cultural resources according to claim 1, characterized in that, The application of the cultural resource big data reference model is specifically to apply the cultural resource big data reference model to online and offline systems. The framework of public cultural digital resources grasps the effective combination of online and offline, and the application is subdivided into the application of library data resources, museum data resources, public cultural cloud data resources, and cultural center data resources.
Citation Information
Patent Citations
Deep theme model-based text image multi-mode retrieval method
CN107609055A
Automatic annotation system for cultural resource data
CN111309933A