Cultural data classification method and device based on digital processing and electronic equipment

By employing multimodal data fusion and dynamic threshold adjustment, the problem of low efficiency in dating and identifying guqin scores has been solved, and the quality and efficiency of automated classification of ancient books have been improved. This approach adapts to different preservation conditions and regional styles, and supports processing speeds down to the second level.

CN120976943APending Publication Date: 2025-11-18SHANDONG POLYTECHNIC COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511103458.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In the field of ancient book digitization, existing technologies are inefficient in dating and identifying ancient zither scores, and are prone to misjudgment due to ink diffusion, paper damage, and carbonization adhesion caused by aging. Traditional methods fail to effectively utilize the temporal characteristics of musical rhythms and adaptive threshold settings, making it difficult to adapt to different preservation conditions and regional styles.

Method used

By employing a multimodal data fusion architecture, the topological features and melody sequence features of musical notation symbols are extracted. Combined with carbonization influence factors and dynamic threshold adjustments, the fifth degree of generation index and jump index are constructed to achieve automated classification of guqin scores.

Benefits of technology

It has achieved a dual improvement in quality and efficiency of ancient genealogy dating technology, enhanced the distinguishability and robustness of dating features, achieved a processing speed of seconds, broken the dependence on manual experience, and provided a standardized technical path for the automated organization of massive amounts of ancient books.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976943A_ABST
    Figure CN120976943A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a cultural data classification method and device based on digital processing and electronic equipment, and the method comprises the steps: extracting the center point coordinates of a target image music score symbol, and extracting the topological features of the music score symbol, the topological features including an adjacent feature and a spacing feature; determining candidate images according to the topological characteristics of the music score symbols; constructing a carbonization influence factor according to the number of target storage years, and updating a determination process of the candidate image based on the carbonization influence factor; constructing a five-degree phase generation index of the candidate image, and extracting a jump index of the candidate image according to a construction result of the five-degree phase generation index; and constructing feature parameters of the candidate images according to the adjacency features of the music score symbols of the candidate images and the jump indexes of the candidate images, and classifying the candidate images according to the feature parameters. According to the invention, the digitalized classification efficiency and classification accuracy of the collected ancient books are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method, apparatus, and electronic device for classifying cultural data based on digital processing. Background Technology

[0002] In the field of ancient book digitization, the dating and identification of guqin scores has long relied on manual experience comparison, which suffers from problems such as low efficiency and significant subjective bias. Traditional automated methods are mostly based on symbolic morphological features for analysis, but the notation symbols of Tang Dynasty tablature and Song Dynasty gongche notation have high topological similarity, making them prone to misjudgment due to ink diffusion and paper damage.

[0003] Existing technologies suffer from three major bottlenecks: First, they focus on a single image modality, neglecting the complementary value of temporal features related to rhythm and phonology; second, they fail to consider the carbonization and adhesion effects caused by the preservation age of ancient books, and the overlapping interference of symbols in aging texts severely reduces the reliability of topological parameters; third, traditional threshold settings rely on static empirical values, making it difficult to adapt to the gradual changes in the characteristics of different preservation states and regional styles of musical scores. These problems limit the engineering application of rapid classification of large quantities of ancient books, and a multi-dimensional, adaptive solution is urgently needed to break through the technological ceiling. Summary of the Invention

[0004] The purpose of this invention is to provide a method, apparatus, and electronic device for classifying cultural data based on digital processing, so as to solve at least one of the problems existing in the prior art.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: A method for classifying cultural data based on digital processing, comprising: Extract the center point coordinates of the musical notation symbols in the target image, and extract the topological features of the musical notation symbols, including adjacency features and spacing features; Candidate images are determined based on the topological features of musical notation symbols; The process of determining carbonization impact factors based on the target preservation years and updating candidate images based on carbonization impact factors; Construct the five-degree begonia index of candidate images, and extract the jump index of candidate images based on the construction results of the five-degree begonia index; Feature parameters of candidate images are constructed based on the adjacency features of musical notation symbols and the jump index of candidate images, and candidate images are classified based on the feature parameters.

[0006] Optionally, the coordinates of the rectangular bounding boxes of the musical scores are extracted based on the pre-trained OCR model, and the coordinates of the center point of the musical scores are calculated using the coordinates of the rectangular bounding boxes. A circular region with the center point of the musical scores as the center and a radius of length m is taken as the adjacent region of the musical scores. The number of center points of the musical scores in the adjacent region is taken as the adjacency density of the musical scores, denoted as Di, where i is the musical score number and i is a positive integer. The average adjacency density of each musical score in the target image is calculated as the adjacency feature of the target image, denoted as Dp.

[0007] Optionally, identical musical notation symbols are grouped, and each group of musical notation symbols is paired up. The distance between each pair of musical notation symbols is calculated using the Euclidean distance formula, and the minimum distance between each pair of musical notation symbols in each group is taken as the spacing feature of that group of musical notation symbols, denoted as Lk, where k is the group number and k is a positive integer.

[0008] Optionally, if Dp is less than or equal to the first quantity threshold and there exists Lk greater than or equal to the first distance threshold, then the target image is determined to be a candidate image; If Dp is greater than or equal to the second quantity threshold and there exists Lk less than or equal to the second distance threshold, then the target image is determined to be a candidate image; If it does not fall into either of the above two categories, the target image will not be classified as a candidate image.

[0009] Optionally, when the target preservation years ns are greater than a preset age threshold, a carbonization impact factor is constructed. The expression for the carbonization impact factor is: ty=exp[-0.01×(ns / 100)]; ty is the carbonization impact factor. The adjacency features are updated to Dp', and Dp' = Dp / ty is set to update the candidate image determination process; The process of determining candidate images is not updated when the target retention period ns is less than or equal to a preset age threshold.

[0010] Optionally, the extracted pitch sequence of the candidate image is denoted as [p1,p2,...,pz], where p1 is the pitch value of the first element in the pitch sequence, p2 is the pitch value of the second element in the pitch sequence, pz is the pitch value of the z-th element in the pitch sequence, and z is the total number of elements in the pitch sequence. Traverse adjacent intervals in the pitch sequence and set |p k+1 -p k |=7 is denoted as a perfect fifth event, and the occurrence frequency of perfect fifth events is counted. The ratio of the occurrence frequency of perfect fifth events to the total number of elements in the pitch sequence is taken as the p-5 Pythagorean index. k+1 p is the pitch value of the (k+1)th element in the pitch sequence. k Let be the pitch value of the k-th element in the pitch sequence, where k∈[1,z-1].

[0011] Optionally, if the fifth degree of co-occurrence index of a candidate image is greater than the first fifth degree of co-occurrence threshold or less than the second fifth degree of co-occurrence threshold, the jump index of the candidate image is extracted; otherwise, the candidate image is removed. The process of extracting the jump index of candidate images is as follows: Traverse adjacent intervals in the pitch sequence and set |p k+1 -p k |≥3 is denoted as a mutation event, and the corresponding |p is set as follows: k+1 -p k | As the mutation value of this mutation event; A jump index is constructed based on the mutation value of each mutation event. The expression for the jump index is as follows: In the formula, tm is the mutation value of the m-th mutation event, and M is the number of mutation events.

[0012] Optionally, the expression for the feature parameters of the candidate image is: C = w1 × (adjacent features / Dmax) + w2 × (1 - J / Jmax); Where C is the feature parameter of the candidate image, w1 is the adjacency feature weight, w2 is the jump weight, w1+w2=1, Dmax is the maximum adjacency threshold, and Jmax is the maximum jump exponential threshold. If the feature parameters of a candidate image are greater than the classification threshold, the candidate image is determined to be a Tang Dynasty tablature; otherwise, it is determined to be a Song Dynasty musical notation.

[0013] According to another aspect of this application, a cultural data classification device based on digital processing is provided, comprising: The topological feature extraction unit is used to extract the center point coordinates of the musical notation symbols in the target image and extract the topological features of the musical notation symbols, including adjacency features and spacing features. The candidate image determination unit is used to determine candidate images based on the topological features of musical notation symbols; The update unit is used to construct a carbonization impact factor based on the target preservation years and to update the candidate image determination process based on the carbonization impact factor. A jump feature construction unit is used to construct the five-degree-of-five co-occurrence index of the candidate image and extract the jump index of the candidate image based on the construction result of the five-degree-of-five co-occurrence index. The classification unit is used to construct feature parameters of candidate images based on the adjacency features of musical notation symbols in the candidate images and the jump index of the candidate images, and to classify the candidate images based on the feature parameters.

[0014] According to another aspect of this application, an electronic device is provided, comprising: One or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the cultural data classification method based on digital processing.

[0015] The beneficial effects of this invention are as follows: Through an innovative multimodal data fusion architecture, it achieves a dual improvement in the quality and efficiency of ancient music score dating technology. It introduces a joint analysis mechanism of symbol topological distribution and musical rhythm sequences, constructing a comprehensive criterion across visual and auditory data: spatial topological features accurately quantify differences in writing and typesetting, while musical rhythm jump index deeply analyzes the evolution of music theory; the synergistic effect of these two significantly enhances the distinguishability of dating features. A unique dynamic compensation algorithm for carbonization influence factors quantitatively solves the false detection of symbols adhering to old texts, enhancing the robustness of aging document analysis. The entire process adopts a lightweight computing strategy, accelerating the process through adjacency density gridding and dynamically adjusting thresholds, achieving second-level processing speeds while ensuring high-precision classification. This solution breaks away from reliance on manual experience, providing a standardized technical path for the automated organization of massive amounts of ancient books, promoting the leap from data storage to intelligent analysis in the digitization of cultural heritage, and possessing scalable value for fields such as ancient music restoration and document preservation. (See attached figures.) To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a flowchart illustrating the cultural data classification method based on digital processing in this embodiment.

[0017] Figure 2 This is a flowchart illustrating the method for determining candidate images in this embodiment.

[0018] Figure 3 This is a flowchart illustrating the method for extracting the jump index in this embodiment.

[0019] Figure 4 This is a schematic diagram of the structure of the cultural data classification device based on digital processing in this embodiment.

[0020] Figure 5 This is a schematic diagram of the electronic device in this embodiment. Detailed Implementation

[0021] To more clearly illustrate the present invention, the following description, in conjunction with preferred embodiments and accompanying drawings, further explains the invention. Similar components in the drawings are indicated by the same reference numerals. Those skilled in the art should understand that the specific description below is illustrative rather than restrictive and should not be construed as limiting the scope of protection of the present invention.

[0022] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0023] Specifically, this embodiment is applied to the automated classification needs of Tang Dynasty tablature and Song Dynasty gongche notation in the digitization project of ancient books in the collection. It solves the problem of low efficiency of manual verification due to the similarity of the notation symbols of the two types of musical scores. It achieves rapid dating by combining symbol topological relationships and musical characteristics.

[0024] Please see Figure 1 As shown, it is a flowchart illustrating the cultural data classification method based on digital processing in this embodiment, including: Step S101: Extract the center point coordinates of the musical notation symbols in the target image and extract the topological features of the musical notation symbols. The target image is a grayscale image of a high-resolution scan of a guqin score, with a ratio of 1:1 to the original.

[0025] For example, in this embodiment, the scanned document can be obtained by a scanner, and the grayscale image of the scanned document can be obtained by OpenCV with a resolution of 600dpi. In this embodiment, the method of obtaining the target image is not specifically limited, and those skilled in the art can set it freely according to their needs.

[0026] Specifically, through standardized symbol localization and coordinate extraction, a standardized geometric feature description system is established, providing a unified data benchmark for subsequent topology analysis. This avoids feature distortion caused by differences in scanning scale or image distortion, ensuring the reliability of the classification model input. It also addresses the shortcomings of traditional visual inspection in terms of insufficient coordinate quantization.

[0027] Please see Figure 2 As shown, the method for determining the candidate image includes: Step S201: Extract the center point coordinates of the musical notation symbols in the target image and extract the adjacency features of the musical notation symbols.

[0028] Specifically, the coordinates of the rectangular bounding boxes of musical scores are extracted based on the pre-trained OCR model. The coordinates of the center point of the musical score are then calculated using the bounding box coordinates. A circular region with the center point of the musical score as the center and a radius of length m is taken as the adjacent region of the musical score. The number of center points of the musical score within the adjacent region is taken as the adjacency density of the musical score, denoted as Di, where i is the musical score number and i is a positive integer. The average adjacency density of each musical score in the target image is calculated as the adjacency feature of the target image, denoted as Dp.

[0029] For example, in this embodiment, m can be set to 71 pixels. This embodiment does not specifically limit the setting of m, and those skilled in the art can set it freely according to their needs.

[0030] Specifically, based on the annular spatial analysis of adjacent regions, the distribution density of musical notation symbols is accurately characterized, and the macroscopic layout differences in notation styles across different historical periods are quantified. Through the key parameter of symbol density, the sparse layout of Tang dynasty scores and the compact arrangement of Song dynasty scores are effectively distinguished, overcoming the limitations of comparing single symbol forms. Simultaneously, the design of the central neighborhood suppresses false detection interference from edge symbols.

[0031] Please continue reading. Figure 2 As shown, the method for determining candidate images further includes: Step S202: Extract the spacing features of the musical notation symbols.

[0032] Specifically, identical musical notation symbols are grouped, and each group of musical notation symbols is paired up. The distance between each pair of musical notation symbols is calculated using the Euclidean distance formula, and the minimum distance between each pair of musical notation symbols in each group is taken as the spacing feature of that group of musical notation symbols, denoted as Lk, where k is the group number and k is a positive integer.

[0033] Specifically, if a set of musical notation symbols has 3 or fewer elements, then that set of musical notation symbols will not be analyzed.

[0034] Specifically, the study focuses on detecting the minimum spacing between reused symbols of the same type to capture the temporal evolution of writer behavior patterns. An intra-group pairing distance calculation strategy is employed to eliminate the interference of isolated symbols on statistical results and enhance sensitivity to core typesetting rules. This feature identifies the stylistic characteristics of Tang dynasty musical scores, which emphasize rhythmic spacing, in stark contrast to the efficient notation requirements of Song dynasty scores.

[0035] Please continue reading. Figure 1 As shown, the cultural data classification method based on digital processing also includes: Step S102: Determine candidate images based on the topological features of the musical notation symbols.

[0036] Specifically, if Dp is less than or equal to the first quantity threshold and there exists Lk greater than or equal to the first distance threshold, then the target image is determined to be a candidate image; If Dp is greater than or equal to the second quantity threshold and there exists Lk less than or equal to the second distance threshold, then the target image is determined to be a candidate image; If it does not fall into either of the above two categories, the target image will not be classified as a candidate image.

[0037] For example, in this embodiment, the first quantity threshold can be set to 2, the second quantity threshold can be set to 4, the first distance threshold can be set to 3540 pixels, and the second distance threshold can be set to 1888 pixels; this embodiment does not specifically limit the above settings, and those skilled in the art can set them freely according to their needs.

[0038] For example, in this embodiment, a bounding box dataset of pre-annotated guqin notation symbols (jianzipu, gongchepu) can be used to train object detection networks such as YOLO or Faster R-CNN. The model outputs the rectangular bounding box coordinates of each symbol, namely the upper left corner (x1, y1) and the lower right corner (x2, y2). (x1+x2) / 2 is used as the x-coordinate of the center point of the musical notation symbol, and (y1+y2) / 2 is used as the y-coordinate of the center point of the musical notation symbol. The number of center points of musical notation symbols in the adjacent region of the musical notation symbol is obtained through Python's scipy.spatial.KDTree. In this embodiment, the above settings are not specifically limited, and those skilled in the art can set them freely according to their needs.

[0039] For example, in this embodiment, musical notation symbols can be grouped using an OCR tool (e.g., all "gong" symbols can be grouped together), and the distance between each pair of musical notation symbols can be calculated using Python + OpenCV. This embodiment does not impose specific limitations on the above settings, and those skilled in the art can freely set them according to their needs.

[0040] Specifically, a dual-threshold criterion is used to combine the static characteristics of symbol distribution density with the dynamic patterns of reuse intervals to achieve coarse-grained screening. Nonlinear decision rules are employed to avoid misclassification caused by fluctuations in a single parameter, initially filtering out anomalous samples that significantly deviate from the target era and reducing the computational load for subsequent refined analysis.

[0041] Please continue reading. Figure 1 As shown, the cultural data classification method based on digital processing also includes: Step S103: Construct a carbonization influence factor based on the target number of years of preservation, and update the candidate image determination process based on the carbonization influence factor, wherein the target is the original guqin score.

[0042] For example, in this embodiment, the target number of years to be stored can be obtained interactively. This embodiment does not specifically limit the method of obtaining the target number of years to be stored, and those skilled in the art can set it freely according to their needs.

[0043] Specifically, when the target preservation years ns are greater than the preset age threshold, a carbonization impact factor is constructed. The expression for the carbonization impact factor is: ty=exp[-0.01×(ns / 100)]; ty is the carbonization impact factor. The adjacency features are updated to Dp', and Dp' = Dp / ty is set to update the candidate image determination process; The process of determining candidate images is not updated when the target retention period ns is less than or equal to a preset age threshold.

[0044] Specifically, in this embodiment, when constructing the carbonization impact factor, the unit of the target preservation years is not considered; only its numerical value is considered.

[0045] For example, in this embodiment, the preset age threshold can be set to 800 years. This embodiment does not specifically limit the setting of the preset age threshold, and those skilled in the art can set it freely according to their needs.

[0046] Specifically, a dynamic compensation mechanism based on the preservation period of documents is introduced, and the physical interference of carbonization on symbol adhesion is quantified through an exponential decay model. This factor adaptively adjusts the adjacency density criterion, eliminating false detections of symbol overlap caused by aging and degradation, and improving the classification robustness of ancient books with long preservation periods. It also addresses the blind spot in traditional methods that ignore the impact of material degradation on geometric features.

[0047] Please continue reading. Figure 1 As shown, the cultural data classification method based on digital processing also includes: Step S104: Construct the five-degree begat index of the candidate image, and extract the jump index of the candidate image based on the construction result of the five-degree begat index.

[0048] Please see Figure 3 As shown, the method for extracting the jump index includes: Step S301: Construct the five-degree co-occurrence index of the candidate image.

[0049] Specifically, the extracted pitch sequence of the candidate image is denoted as [p1,p2,...,pz], where p1 is the pitch value of the first element in the pitch sequence, p2 is the pitch value of the second element in the pitch sequence, pz is the pitch value of the z-th element in the pitch sequence, and z is the total number of elements in the pitch sequence. Traverse adjacent intervals in the pitch sequence and set |p k+1 -p k|=7 is denoted as a perfect fifth event, and the occurrence frequency of perfect fifth events is counted. The ratio of the occurrence frequency of perfect fifth events to the total number of elements in the pitch sequence is taken as the p-5 Pythagorean index. k+1 p is the pitch value of the (k+1)th element in the pitch sequence. k Let be the pitch value of the k-th element in the pitch sequence, where k∈[1,z-1].

[0050] For example, in this embodiment, musical notation can be parsed into pitch sequences using a MIDI conversion tool and encoded into numbers according to the twelve-tone equal temperament (C=0, C#=1, D=2,...,B=11), such as the pitch sequence of the Tang Dynasty excerpt "Jieshi Tune: Youlan" as [0, 7, 2, 9, 5, 0, 7]. This embodiment does not impose specific limitations on the above settings, and those skilled in the art can freely set them according to their needs.

[0051] Specifically, by analyzing the frequency spectrum patterns of pitch sequences, the characteristics of the historical transmission of music theory are extracted, and the degree of fifth correlation between adjacent pitch transitions is quantified, providing a phonological basis for dating. This overcomes the limitations of simple symbolic analysis in discerning the connotations of the phonological system and enhances the multidimensional interpretive capabilities of the classification system.

[0052] Please continue reading. Figure 3 As shown, the method for extracting the jump index further includes: Step S302: Extract the jump index of the candidate image based on the construction result of the five-degree co-occurrence index.

[0053] Specifically, when the fifth degree of co-occurrence index of a candidate image is greater than the first fifth degree of co-occurrence threshold or less than the second fifth degree of co-occurrence threshold, the jump index of the candidate image is extracted; otherwise, the candidate image is removed. The process of extracting the jump index of candidate images is as follows: Traverse adjacent intervals in the pitch sequence and set |p k+1 -p k |≥3 is denoted as a mutation event, and the corresponding |p is set as follows: k+1 -p k | As the mutation value of this mutation event; A jump index is constructed based on the mutation value of each mutation event. The expression for the jump index is as follows: In the formula, tm is the mutation value of the m-th mutation event, and M is the number of mutation events.

[0054] For example, in this embodiment, the first fifth degree of generation threshold can be set to 0.6, and the second fifth degree of generation threshold can be set to 0.35. This embodiment does not specifically limit the above settings, and those skilled in the art can set them freely according to their needs.

[0055] Specifically, based on the detection of abrupt changes in discontinuous intervals, the evolutionary differences in musical notation systems are captured, and high-frequency jump features complement the symbol typography features, thereby enhancing the ability to identify antique forged scores.

[0056] Please continue reading. Figure 1 As shown, the method for extracting the jump index further includes: Step S105: Construct feature parameters of candidate images based on the adjacency features of musical notation symbols in candidate images and the jump index of candidate images, and classify candidate images based on feature parameters.

[0057] Specifically, the expression for the feature parameters of the candidate image is: C = w1 × (adjacent features / Dmax) + w2 × (1 - J / Jmax); Where C is the feature parameter of the candidate image, w1 is the adjacency feature weight, w2 is the jump weight, w1+w2=1, Dmax is the maximum adjacency threshold, and Jmax is the maximum jump exponential threshold. If the feature parameters of a candidate image are greater than the classification threshold, the candidate image is determined to be a Tang Dynasty tablature; otherwise, it is determined to be a Song Dynasty musical notation.

[0058] For example, in this embodiment, the critical feature weight can be set to 0.7, the jump weight can be set to 0.3, the maximum adjacency threshold can be set to 5, and the maximum jump index threshold can be set to 15. This embodiment does not specifically limit the above settings, and those skilled in the art can set them freely according to their needs.

[0059] Specifically, based on the detection of abrupt changes in discontinuous intervals, the evolutionary differences of musical notation systems are captured. High-frequency jump features and symbol typography features complement each other, enhancing the ability to identify antique forged scores. A weighted fusion strategy enables comprehensive decision-making based on cross-modal evidence.

[0060] Please see Figure 4 As shown, the cultural data classification device based on digital processing includes: The topological feature extraction unit is used to extract the center point coordinates of the musical notation symbols in the target image and extract the topological features of the musical notation symbols, including adjacency features and spacing features. The candidate image determination unit is used to determine candidate images based on the topological features of musical notation symbols; The update unit is used to construct a carbonization impact factor based on the target preservation years and to update the candidate image determination process based on the carbonization impact factor. A jump feature construction unit is used to construct the five-degree-of-five co-occurrence index of the candidate image and extract the jump index of the candidate image based on the construction result of the five-degree-of-five co-occurrence index. The classification unit is used to construct feature parameters of candidate images based on the adjacency features of musical notation symbols in the candidate images and the jump index of the candidate images, and to classify the candidate images based on the feature parameters.

[0061] The cultural data classification device based on digital processing provided in this application can execute the cultural data classification method based on digital processing provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of the execution method.

[0062] Please see Figure 5 As shown, it is a structural schematic diagram of an electronic device in this embodiment. The electronic device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (such as vehicle navigation terminals), wearable electronic devices, etc., as well as fixed terminals such as digital TVs, desktop computers, smart home devices, etc. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0063] like Figure 5 As shown, the electronic device includes: a processor 501, a memory 502, a communication interface 503, and a system bus 504. The processor includes at least one of a central processing unit (CPU), a graphics processing unit (GPU), or a field-programmable gate array (FPGA), configured to call computer programs and data stored in the memory and generate control instructions; the memory includes random access memory (RAM) and / or non-volatile memory (NVM), the NVM including flash memory, solid-state drive (SSD), or a combination thereof, used to store computer programs, process intermediate data, and historical data sets; the communication interface includes a wired communication module and a wireless communication module, the wired communication module supporting Ethernet or RS-485 protocols for connecting to sensor networks; the wireless communication module supporting LoRa, 5G, or satellite communication protocols for transmitting processing results to a remote server; the system bus adopts a PCI Express or AXI bus architecture to achieve high-speed data interaction and clock synchronization between the processor, memory, and communication interface.

[0064] This embodiment also provides a computer-readable storage medium that physically stores computer-executable instructions. When the instructions are transmitted to the processing unit via an integrated circuit substrate, they are encapsulated and processed through the data channel of the bus system and then solidified into the non-volatile storage area of ​​the storage module. The executable instructions are configured to implement the complete technical solution of the cultural data classification method based on digital processing when executed by the processor.

[0065] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. For those skilled in the art, other variations or modifications can be made based on the above description. It is impossible to exhaustively list all the implementation methods here. All obvious variations or modifications derived from the technical solutions of the present invention are still within the protection scope of the present invention.

Claims

1. A method for classifying cultural data based on digital processing, characterized in that, include: Extract the center point coordinates of the musical notation symbols in the target image, and extract the topological features of the musical notation symbols, including adjacency features and spacing features; Candidate images are determined based on the topological features of musical notation symbols; The process of determining carbonization impact factors based on the target preservation years and updating candidate images based on carbonization impact factors; Construct the five-degree begonia index of candidate images, and extract the jump index of candidate images based on the construction results of the five-degree begonia index; Feature parameters of candidate images are constructed based on the adjacency features of musical notation symbols and the jump index of candidate images, and candidate images are classified based on the feature parameters.

2. The cultural data classification method based on digital processing according to claim 1, characterized in that, The coordinates of the rectangular bounding boxes of musical scores are extracted based on the pre-trained OCR model. The coordinates of the center point of the musical score are then calculated using the coordinates of the rectangular bounding boxes. A circular region with the center point of the musical score as the center and a radius of length m is taken as the adjacent region of the musical score. The number of center points of the musical score within the adjacent region is taken as the adjacency density of the musical score, denoted as Di, where i is the musical score number and i is a positive integer. The average adjacency density of each musical score in the target image is calculated as the adjacency feature of the target image, denoted as Dp.

3. The cultural data classification method based on digital processing according to claim 2, characterized in that, Group identical musical notation symbols and pair them together. Calculate the distance between each pair of musical notation symbols using the Euclidean distance formula. Take the minimum distance between each pair of musical notation symbols in each group as the spacing characteristic of that group, denoted as Lk, where k is the group number and k is a positive integer.

4. The cultural data classification method based on digital processing according to claim 3, characterized in that, If Dp is less than or equal to the first quantity threshold and there exists Lk greater than or equal to the first distance threshold, then the target image is determined to be a candidate image; If Dp is greater than or equal to the second quantity threshold and there exists Lk less than or equal to the second distance threshold, then the target image is determined to be a candidate image; If it does not fall into either of the above two categories, the target image will not be classified as a candidate image.

5. The cultural data classification method based on digital processing according to claim 4, characterized in that, When the target preservation years ns are greater than the preset age threshold, a carbonization influence factor is constructed. The expression for the carbonization influence factor is: ty=exp[-0.01×(ns / 100)]; ty is the carbonization influence factor; The adjacency features are updated to Dp', and Dp' = Dp / ty is set to update the candidate image determination process; The process of determining candidate images is not updated when the target retention period ns is less than or equal to a preset age threshold.

6. The cultural data classification method based on digital processing according to claim 5, characterized in that, The extracted pitch sequence of the candidate image is denoted as [p1,p2,...,pz], where p1 is the pitch value of the first element in the pitch sequence, p2 is the pitch value of the second element in the pitch sequence, pz is the pitch value of the z-th element in the pitch sequence, and z is the total number of elements in the pitch sequence. Traverse adjacent intervals in the pitch sequence and set |p k+1 -p k |=7 is denoted as a perfect fifth event, and the occurrence frequency of perfect fifth events is counted. The ratio of the occurrence frequency of perfect fifth events to the total number of elements in the pitch sequence is taken as the p-5 Pythagorean index. k+1 p is the pitch value of the (k+1)th element in the pitch sequence. k Let be the pitch value of the k-th element in the pitch sequence, where k∈[1,z-1].

7. The cultural data classification method based on digital processing according to claim 6, characterized in that, If the fifth degree of commensuration index of a candidate image is greater than the first fifth degree of commensuration threshold or less than the second fifth degree of commensuration threshold, the jump index of the candidate image is extracted; otherwise, the candidate image is removed. The process of extracting the jump index of candidate images is as follows: Traverse adjacent intervals in the pitch sequence and set |p k+1 -p k |≥3 is denoted as a mutation event, and the corresponding |p is set as follows: k+1 -p k | As the mutation value of this mutation event; A jump index is constructed based on the mutation value of each mutation event. The expression for the jump index is as follows: In the formula, tm is the mutation value of the m-th mutation event, and M is the number of mutation events.

8. The cultural data classification method based on digital processing according to claim 7, characterized in that, The expression for the feature parameters of the candidate image is: C = w1 × (adjacent features / Dmax) + w2 × (1 - J / Jmax); Where C is the feature parameter of the candidate image, w1 is the adjacency feature weight, w2 is the jump weight, w1+w2=1, Dmax is the maximum adjacency threshold, and Jmax is the maximum jump exponential threshold. If the feature parameters of a candidate image are greater than the classification threshold, the candidate image is determined to be a Tang Dynasty tablature; otherwise, it is determined to be a Song Dynasty musical notation.

9. A cultural data classification device based on digital processing, applied to the cultural data classification method based on digital processing as described in any one of claims 1-8, characterized in that, include: The topological feature extraction unit is used to extract the center point coordinates of the musical notation symbols in the target image and extract the topological features of the musical notation symbols, including adjacency features and spacing features. The candidate image determination unit is used to determine candidate images based on the topological features of musical notation symbols; The update unit is used to construct a carbonization impact factor based on the target preservation years and to update the candidate image determination process based on the carbonization impact factor. A jump feature construction unit is used to construct the five-degree-of-five co-occurrence index of the candidate image and extract the jump index of the candidate image based on the construction result of the five-degree-of-five co-occurrence index. The classification unit is used to construct feature parameters of candidate images based on the adjacency features of musical notation symbols in the candidate images and the jump index of the candidate images, and to classify the candidate images based on the feature parameters.

10. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the cultural data classification method based on digital processing as described in any one of claims 1-8.