Digital content adaptive generation method and system based on scene semantic understanding
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIAMEN DUOXIANG ANIMATION CO LTD
- Filing Date
- 2026-05-15
- Publication Date
- 2026-08-07
AI Technical Summary
[0004]本发明的目的在于提供基于场景语义理解的数字内容自适应生成方法及系统,旨在解决背景技术中所提到的问题
[0056] This invention constructs multiple data chains by alternately inserting mirror data units with different sequence biases and setting breakpoints within the chains. This transforms the data from a mirror set into a multi-chain sequence structure. Each data unit is assigned a specific chain affiliation and position within the chain, enabling the system to express the adjacency and local sequence relationships between data units. This provides intra-chain context information for subsequent analysis. Through the alternating insertion mechanism, data units with different biases are interleaved within the chains, creating comparable arrangement relationships between chains. The breakpoints prevent the chains from becoming continuous, indivisible sequences, instead providing insertable and reconfigurable local structural units. This transforms the originally planar data set into a data structure with multiple sequence paths, providing a structural foundation for subsequent cross-chain operations.
Smart Images

Figure CN122220345B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method and system for adaptive generation of digital content based on scene semantic understanding. Background Technology
[0002] With the development of artificial intelligence and multimedia technologies, digital content generation methods based on scene semantic understanding are gradually being applied to fields such as digital tourism, exhibitions, and immersive spaces. Existing technologies typically collect images, videos, or sensor data to identify and analyze objects, environments, and user behaviors within a scene. They then utilize computer vision and semantic modeling techniques to extract semantic information from the scene. Based on this, and combined with preset rules or model algorithms, corresponding digital content templates or generation models are invoked to generate and display digital content. Simultaneously, through the coordinated control of display terminals, projection equipment, and intelligent control systems, the generated content is presented in a specific space, achieving a certain degree of interactive experience and content linkage effects.
[0003] However, in practical applications, the semantic understanding results of existing technologies are usually associated with digital content generation through fixed mapping relationships or limited rules, which may lack the ability to deeply model the semantics of complex scenarios. For example, in immersive exhibition halls or CAVE spaces, when visitors move or gather in a certain exhibit area, although the system can detect changes in the location or number of people, because the semantic understanding remains at a superficial level of information such as location or number, it may not be able to further identify the user's interests or behavioral intentions. As a result, the generated digital content is still mainly based on preset logic and may not be able to dynamically adjust to the specific content that the audience is interested in. This leads to a mismatch between the displayed content and the user's actual focus, affecting the accuracy and immersion of the overall interactive experience. Summary of the Invention
[0004] The purpose of this invention is to provide a method and system for adaptive generation of digital content based on scene semantic understanding, in order to solve the problems mentioned in the background art.
[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:
[0006] Firstly, a digital content adaptive generation method based on scene semantic understanding, the method comprising:
[0007] Obtain the input dataset, copy each data unit in the input dataset to generate at least two mirror data units, and assign different sequence bias labels to each mirror data unit to obtain the mirror dataset;
[0008] Based on the mirror dataset, mirror data units with different sequence bias identifiers are alternately inserted to construct multiple data chains, and breakpoint positions are set in each data chain to obtain interleaved chain data;
[0009] Based on the interleaved chain data, data units from each data chain are introduced at each breakpoint, and a multi-chain interleaved network is constructed according to the interval of the breakpoint to obtain the interleaved structure data.
[0010] Based on the interwoven structure data, the frequency of the same data unit appearing in different positions in each data chain and the cross-chain distribution range are counted to identify the concentration of data units and obtain the structural aggregation quantity.
[0011] Based on the interleaved structure data, each data chain is divided into several segments according to the variation range of the data units, and the data units within the segments are folded to obtain folded segment data;
[0012] Based on the folded segment data, the source data chain of the data unit is traced step by step from the end of each segment. The number of path forks during the tracing process is recorded, the degree of data change and offset is identified, and the sequence offset is obtained.
[0013] By mapping structural aggregation and sequence offset to aggregation level intervals and offset level intervals, and combining different level intervals to establish rule encoding, rule data is obtained.
[0014] Based on the rule data, the preset content resources are divided into multiple content units, and the content units are coded and matched for different rules to obtain the target content data.
[0015] Furthermore, based on the mirror dataset, mirror data units with different sequence bias identifiers are alternately inserted to construct multiple data chains, and breakpoints are set in each data chain to obtain interleaved chain data, including:
[0016] Based on the mirror dataset, each mirror data unit is divided into multiple bias unit groups according to the sequence bias identifier, and the order of each mirror data unit in the bias unit group is recorded to obtain the bias grouped data.
[0017] Based on the bias grouped data, the order of the mirrored data units within the same bias unit group is maintained, and the mirrored data units are written sequentially into the sorting queue to obtain the alternating queue data.
[0018] Based on the alternating queue data, the mirror data units in the permutation queue are assigned to different chain sequences, and the chain number and position within the chain of each mirror data unit are recorded to obtain the initial chain data;
[0019] Based on the initial chain data, breakpoint identifiers are written into each chain sequence, and the breakpoint identifiers are associated with the chain number, the position within the chain, and the position of adjacent mirror data units to obtain interleaved chain data.
[0020] Furthermore, based on the interleaved chain data, data units from each data chain are introduced at each breakpoint, and a multi-chain interleaved network is constructed according to the intervals of the breakpoints to obtain interleaved structure data, including:
[0021] Based on the interleaved chain data, extract the breakpoint identifiers, chain numbers, and intra-chain positions in each data chain, calculate the interval positions between adjacent breakpoint identifiers, and obtain the breakpoint interval data.
[0022] Based on the breakpoint interval data, select the mirror data unit corresponding to the interval position according to the chain number, establish the position mapping relationship between the selected mirror data unit and the breakpoint identifier, and obtain the introduced unit data;
[0023] Based on the introduced unit data, the selected mirror data unit is written to the position of the breakpoint identifier, and the writing position is adjusted forward and backward according to the breakpoint interval data to obtain the embedded chain data.
[0024] Based on the embedded chain data, the cross-chain connection relationships formed by the breakpoint embedding in each data chain are recorded, and the cross-chain connection relationships are arranged and combined into a multi-chain intertwined relationship structure to obtain the intertwined structure data.
[0025] Furthermore, based on the interwoven structure data, the frequency of occurrence of the same data unit in different positions and the cross-chain distribution range in each data chain are statistically analyzed to identify the concentration of data units and obtain the structural aggregation quantity, including:
[0026] Based on the interleaved structure data, the number of times the same data unit appears in all data chains is counted to identify the density of the data unit in the multi-chain interleaved relationship structure and obtain the occurrence intensity term.
[0027] Based on the interleaved structure data, the span value between data chain numbers containing the same data unit is calculated, the cross-chain convergence degree of the same data unit is identified, and the distributed convergence term is obtained;
[0028] Based on the interleaved structure data, the difference between the maximum and minimum position of the data unit in the single data chain is calculated to identify the degree of positional clustering of the data unit in the single chain and obtain the clustering term in the chain.
[0029] Based on the interleaved structure data, calculate the position center value of the same data unit in each data chain, identify the degree of position consistency of the data unit between different data chains, and obtain the position consistency item;
[0030] By fusing the intensity term, distribution convergence term, intra-chain aggregation term, and positional consistency term, the concentration of data units in each data chain is identified, and the structural aggregation quantity is obtained.
[0031] Furthermore, based on the interleaved structure data, each data chain is divided into several segments according to the variation range of the data units, and the data units within each segment are folded to obtain folded segment data, including:
[0032] Based on the interleaved structure data, the intra-chain positional difference of adjacent data units in each data chain is calculated bit by bit, and the intra-chain positional difference is associated with the change of the identifier to obtain the change record data;
[0033] Based on the change record data, according to the change of identifiers between continuous data units, data units with the same change trend are divided into the same segment to obtain segmentation data;
[0034] Based on the segmented data, the data units within the same segment are merged and replaced with the original data units according to their arrangement order. At the same time, the position range of the merged data units is recorded to obtain the folded unit data.
[0035] Based on the folded unit data, each folded unit is rearranged according to its data chain and segment identifier, establishing a correspondence between the folded unit and the original data chain to obtain the folded segment data.
[0036] Furthermore, based on the folded segment data, the source data chain of the data unit is traced step by step from the end of each segment. The number of path forks during the tracing process is recorded, the degree of data shift is identified, and the sequence offset is obtained, including:
[0037] Based on the folded segment data, the number of source forks in each folded segment is counted, the degree of diffusion of the folded segment in the source tracing process is identified, and the fork extension term is obtained;
[0038] Based on the folded segment data, calculate the intra-chain positional change of each folded segment, identify the degree of positional jump of the folded segment during the backtracking process, and obtain the positional jump term;
[0039] Based on the folded segment data, the activity level of cross-chain changes in the folded segment during the backtracking process is identified to obtain the inter-chain switching item;
[0040] By fusing the fork extension term, the positional shift term, and the inter-chain switching term, the overall change and offset of the folded segment during the backtracking process are identified, and the sequence offset is obtained.
[0041] Furthermore, by mapping structural aggregation and sequence offset to aggregation level intervals and offset level intervals, and combining different level intervals to establish rule encoding, rule data is obtained, including:
[0042] By dividing the structural aggregation quantity into multiple aggregation level intervals according to preset aggregation values, and assigning an aggregation interval identifier to each aggregation level interval, aggregation interval data is obtained.
[0043] By dividing the sequence offset into multiple offset level intervals according to preset offset values and assigning an offset interval identifier to each offset level interval, offset interval data is obtained.
[0044] By merging each aggregation level interval with each offset level interval into a combined interval, and establishing a combined identifier for each combined interval, the interval combined data is obtained;
[0045] Based on the interval combination data, the combination identifier is converted into a rule code according to the preset encoding rules, and a mapping relationship between the rule code and the combination interval is established to obtain the rule data.
[0046] Secondly, a digital content adaptive generation system based on scene semantic understanding, the system comprising:
[0047] The mirroring module is used to obtain the input dataset, copy each data unit in the input dataset to generate at least two mirror data units, and assign different sequence bias labels to each mirror data unit to obtain the mirror dataset;
[0048] The interleaving module is used to alternately insert mirror data units with different sequence bias identifiers into multiple data chains based on the mirror dataset, and set breakpoint positions in each data chain to obtain interleaved chain data;
[0049] The interleaving module is used to introduce data units from each data chain at each breakpoint position based on the interleaved chain data, and construct a multi-chain interleaving network according to the interval of the breakpoint positions to obtain interleaved structure data.
[0050] The aggregation module is used to count the number of times the same data unit appears in different positions in each data chain and the cross-chain distribution range based on the interwoven structure data, identify the concentration of data units, and obtain the structure aggregation quantity.
[0051] The folding module is used to divide each data chain into several segments according to the variation range of data units based on the interleaved structure data, and fold the data units within the segments to obtain folded segment data;
[0052] The offset module is used to trace the source data chain of the data unit step by step from the end of each segment based on the folded segment data, record the number of path forks during the tracing process, identify the degree of data change and offset, and obtain the sequence offset.
[0053] The rules module is used to map structural aggregation quantities and sequence offsets to aggregation level intervals and offset level intervals, combine different level intervals and establish rule codes to obtain rule data;
[0054] The matching module is used to split the preset content resources into multiple content units according to the rule data, and to encode the matching content units for different rules to obtain the target content data.
[0055] The above-described solution of the present invention has at least the following beneficial effects:
[0056] This invention constructs multiple data chains by alternately inserting mirror data units with different sequence biases and setting breakpoints within the chains. This transforms the data from a mirror set into a multi-chain sequence structure. Each data unit is assigned a specific chain affiliation and position within the chain, enabling the system to express the adjacency and local sequence relationships between data units. This provides intra-chain context information for subsequent analysis. Through the alternating insertion mechanism, data units with different biases are interleaved within the chains, creating comparable arrangement relationships between chains. The breakpoints prevent the chains from becoming continuous, indivisible sequences, instead providing insertable and reconfigurable local structural units. This transforms the originally planar data set into a data structure with multiple sequence paths, providing a structural foundation for subsequent cross-chain operations.
[0057] This invention introduces data units from each data chain at breakpoints and constructs a multi-chain interwoven network based on breakpoint intervals. This establishes connections between previously independent data chains. The data chains are not mutually independent linear structures, but rather form a network structure with cross-chain connections through breakpoint embedding. In this network, a data unit not only exists in its own chain but may also form associations with other chains through breakpoint introductions. This expands the relationships between data from intra-chain relationships to cross-chain relationships. This interwoven structure can express the distribution and mutual reference relationships of data units across different chains, expanding the internal data organization of the system from a linear structure to a network structure. It also makes cross-chain connections positionally relevant, and the expression of data associations is no longer limited to a single sequence but possesses multi-dimensional structural characteristics.
[0058] This invention obtains structural aggregation quantity by statistically analyzing the frequency of occurrence of the same data unit in different positions and the cross-chain distribution range within each data chain of an interwoven structure. It transforms the distribution information in the interwoven network into quantifiable statistical results, realizing the mapping from structural relationships to numerical features. That is, the distribution information, which originally existed in the form of position, chain number, etc., is transformed into a concentration index. By statistically analyzing the frequency of occurrence, it reflects the repetitive distribution of a certain data unit in the overall structure. By statistically analyzing the cross-chain range, it reflects the diffusion or convergence of the data unit between different chains. This compresses the complex relationship of the interwoven structure into a unified aggregation quantity representation, allowing different data units to be compared at the same scale. Subsequent processing can be based on numerical indicators for classification, grading, or mapping without directly manipulating the complex chain or network structure.
[0059] This invention achieves local structural compression of chained data by dividing each data chain into several segments according to the variation range of data units and then folding the data units within each segment. By dividing the data chain into segments based on the variation range, data units with similar variation trends are grouped into the same segment. Folding the data units within each segment merges multiple original data units into a single representation unit while preserving their positional range information. This transforms the originally long and detailed data chain into a compressed structure composed of several segments. The folded segment data not only reduces the number of data units but also retains segment-level structural information, allowing the system to perform analysis at the segment level in subsequent processing without having to process each original data unit individually, thus avoiding changes to the granularity and structural hierarchy of data processing.
[0060] This invention traces the source data chain of data units step by step from the end of each segment, and records the number of path forks during the tracing process to obtain the sequence offset. By recording the number of forks in the path, it reflects the degree of path dispersion experienced by the segment during its formation. By observing the chain changes during the tracing process, it reflects the migration path of data between different chains, transforming the data change path from an implicit structural relationship into an explicit numerical indicator. In subsequent steps, the system can simultaneously utilize the current structural state of the data and its formation path information, introducing the expression of time or evolution dimensions into the data representation, and quantifying and standardizing the path information. This allows complex backtracking relationships to be uniformly encoded and further utilized. Attached Figure Description
[0061] Figure 1 This is a flowchart of a digital content adaptive generation method based on scene semantic understanding provided in an embodiment of the present invention. Detailed Implementation
[0062] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0063] like Figure 1 As shown, embodiments of the present invention propose an adaptive digital content generation method based on scene semantic understanding, the method comprising:
[0064] Obtain the input dataset, copy each data unit in the input dataset to generate at least two mirror data units, and assign different sequence bias labels to each mirror data unit to obtain the mirror dataset;
[0065] Based on the mirror dataset, mirror data units with different sequence bias identifiers are alternately inserted to construct multiple data chains, and breakpoint positions are set in each data chain to obtain interleaved chain data;
[0066] Based on the interleaved chain data, data units from each data chain are introduced at each breakpoint, and a multi-chain interleaved network is constructed according to the interval of the breakpoint to obtain the interleaved structure data.
[0067] Based on the interwoven structure data, the frequency of the same data unit appearing in different positions in each data chain and the cross-chain distribution range are counted to identify the concentration of data units and obtain the structural aggregation quantity.
[0068] Based on the interleaved structure data, each data chain is divided into several segments according to the variation range of the data units, and the data units within the segments are folded to obtain folded segment data;
[0069] Based on the folded segment data, the source data chain of the data unit is traced step by step from the end of each segment. The number of path forks during the tracing process is recorded, the degree of data change and offset is identified, and the sequence offset is obtained.
[0070] By mapping structural aggregation and sequence offset to aggregation level intervals and offset level intervals, and combining different level intervals to establish rule encoding, rule data is obtained.
[0071] Based on the rule data, the preset content resources are divided into multiple content units, and the content units are coded and matched for different rules to obtain the target content data.
[0072] In this embodiment of the invention, an input dataset is obtained, and each data unit in the input dataset is copied to generate at least two mirror data units. Each mirror data unit is assigned a different sequence bias identifier to obtain a mirror dataset, transforming a single-instance format into a multi-instance format from the same source, providing a basis for data differentiation in subsequent operations. Based on the mirror dataset, mirror data units with different sequence bias identifiers are alternately inserted to construct multiple data chains, and breakpoints are set in each data chain to obtain interleaved chain data. Alternating insertion ensures that data units with different bias identifiers form a comparable and interleaved distribution relationship within the same chain, providing a structural interface for subsequent interleaved network construction. Based on the interleaved chain data, data units from each data chain are introduced at each breakpoint, and a multi-chain interleaved network is constructed based on the intervals between breakpoints, resulting in interleaved structure data. This transforms the parallel chain organization into an interleaved structure, providing a structural foundation for subsequent analysis and processing. Based on the interleaved structure data, the frequency of occurrence and cross-chain distribution range of the same data units in different positions in each data chain are statistically analyzed to identify the concentration of data units, obtaining structural aggregation. This converts the complex distribution relationship in the interleaved network into structural statistical results, providing numerical basis for subsequent operations.
[0073] Based on the interwoven structure data, each data chain is divided into several segments according to the variation range of data units. The data units within each segment are then folded to obtain folded segment data. This reconstructs the fine-grained chain data into a data representation, eliminating the need for subsequent source tracing to backtrack to each original unit individually. Based on the folded segment data, the source data chain of each data unit is traced step by step from the end of each segment. The number of path forks during the tracing process is recorded, the degree of data change and offset is identified, and the sequence offset is obtained. The path evolution information is transformed into offset results, providing path offset status characteristics for subsequent processing. By mapping the structural aggregation quantity and sequence offset to aggregation level intervals and offset level intervals, different level intervals are combined to establish rule codes, resulting in rule data. This changes the way front-end data analysis results are directly connected to content resources, making the intermediate rule layer the interface for subsequent content matching. Based on the rule data, the preset content resources are divided into multiple content units, and different rule codes are used to match the content units to obtain target content data. The mapping relationship between rule codes and content units enables rule-oriented matching calls, achieving a closed-loop connection from structured data processing to target data output.
[0074] The process includes obtaining the input dataset, copying each data unit in the input dataset to generate at least two mirror data units, and assigning different sequence bias labels to each mirror data unit to obtain the mirror dataset. Specifically, this includes:
[0075] The input dataset can be a structured data set that has undergone semantic processing. Each data unit contains at least a unique identifier field and several data fields representing its state or attributes. The system iterates through each data unit in the dataset and generates multiple mirror copies for each original data unit. The copying process can be achieved by allocating a new storage address or a new instance reference to the original data unit while retaining its original field values. The system writes a sequence bias identifier to each mirror data unit. This sequence bias identifier can be generated using a preset identifier set, such as incremental numbering, predefined labels, or encoding methods, so that different mirror data units have logically distinguishable sequence attributes. The system also records the mapping relationship between the mirror data unit and the original data unit for traceability or association in subsequent processing. The system summarizes all mirror data units to form a mirror dataset. This mirror dataset maintains consistency with the original input in terms of data content, but its structure has been expanded from a single data instance to a collection of multiple instances.
[0076] Specifically, based on the rule data, the preset content resources are divided into multiple content units, and different rule codes are used to match the content units to obtain the target content data, which includes:
[0077] The system reads a pre-stored content resource library. This content resource can be in various data formats for output, such as text fragments, voice data, graphical control parameters, animation instructions, or other executable content data. The system parses the content resource, dividing it into multiple independent content units according to preset splitting rules. These divisions are based on internal structural boundaries, logical paragraphs, time segments, or functional modules, enabling each content unit to be independently invoked and combined. The system establishes a corresponding unit identifier for each content unit and records its attribute information and applicable rule encoding range in the resource management table. The system reads the rule data generated in the preceding steps and parses the rule codes one by one. For each rule code, the system retrieves a set of content units matching that code from the content resource management table. The matching process can be based on complete matching, interval matching, or mapping relationship matching. When multiple content units correspond to the same rule code, the system can sort and combine these units according to preset priority order, combination strategy, or sequence rules. The system then splices or structurally combines the selected content units according to a predetermined output order to generate the corresponding target content data. During the assembly process, the connection relationship between content units is processed, such as inserting connection identifiers, adjusting the output order, or unifying the data format. The system then outputs the generated target content data to downstream modules or execution units, completing the rule-based data-driven content generation process.
[0078] In a preferred embodiment of the present invention, based on the mirror dataset, mirror data units with different sequence offset identifiers are alternately inserted to construct multiple data chains, and breakpoint positions are set in each data chain to obtain interleaved chain data, including:
[0079] Based on the mirror dataset, each mirror data unit is divided into multiple bias unit groups according to the sequence bias identifier, and the order of each mirror data unit in the bias unit group is recorded to obtain the bias grouped data.
[0080] Based on the bias grouped data, the order of the mirrored data units within the same bias unit group is maintained, and the mirrored data units are written sequentially into the sorting queue to obtain the alternating queue data.
[0081] Based on the alternating queue data, the mirror data units in the permutation queue are assigned to different chain sequences, and the chain number and position within the chain of each mirror data unit are recorded to obtain the initial chain data;
[0082] Based on the initial chain data, breakpoint identifiers are written into each chain sequence, and the breakpoint identifiers are associated with the chain number, the position within the chain, and the position of adjacent mirror data units to obtain interleaved chain data.
[0083] In this embodiment of the invention, based on the mirror dataset, each mirror data unit is divided into multiple bias unit groups according to the sequence bias identifier, and the order of each mirror data unit in the bias unit group is recorded to obtain bias grouping data. This avoids losing the original relative order relationship of the mirror data units during the alternation process and provides grouping and order information for subsequent data chain construction. Based on the bias grouping data, the order of the mirror data units within the same bias unit group is maintained, and the mirror data units are sequentially written into the arrangement queue to obtain alternation queue data, so that data units with different bias identifiers are processed in sequence. The chain-like structure has basic arrangement conditions. Based on the alternating queue data, the mirror data units in the arrangement queue are assigned to different chain sequences. The chain number and intra-chain position of each mirror data unit are recorded to obtain the initial chain data, which provides the positioning basis for setting breakpoints in each chain and establishing the relationship between breakpoints and adjacent units. Based on the initial chain data, breakpoint identifiers are written into each chain sequence, and the breakpoint identifiers are associated with the chain number, intra-chain position, and position of adjacent mirror data units to obtain the interleaved chain data. The position coordinates of the breakpoints and the local adjacency boundaries are clarified, providing a structural reference for subsequent operations.
[0084] Specifically, based on the bias grouping data, the mirrored data units within the same bias unit group are kept in the same order and sequentially written into the permutation queue to obtain the alternating queue data, which includes:
[0085] The bias grouping data includes at least multiple bias unit groups and the order of each mirror data unit within its corresponding bias unit group. When constructing the queue, the system extracts mirror data units from each bias unit group in an orderly manner, maintaining the order within each group. An independent read pointer is established for each bias unit group, initially pointing to the starting position of the order within that group. Following a preset inter-group rotation order, the system sequentially selects the mirror data unit pointed to by the current read pointer from different bias unit groups and writes it to the tail of the queue. After one write operation, the system moves the read pointer of the corresponding bias unit group one position forward to prepare for the next round of extraction. If multiple bias unit groups contain unprocessed mirror data units, the system continues to execute data writing according to the aforementioned rotation strategy, resulting in an alternating distribution of mirror data units in the queue across different bias unit groups. When all mirror data units in a certain bias unit group have been written, the system automatically skips that bias unit group and continues rotating the writing process for the remaining bias unit groups until all mirror data units have been written to the queue. The system records the queue position index for each mirrored data unit that enters the permutation queue, while retaining its original bias unit group identifier and the permutation order information within the group. The final permutation queue is a linear data structure in which mirrored data units at adjacent positions usually come from different bias unit groups, but the relative order of data units within each bias unit group in the queue remains consistent with the grouping stage, resulting in alternating queue data.
[0086] Specifically, based on the initial chain data, breakpoint identifiers are written into each chain sequence, and these breakpoint identifiers are associated with the chain number, the intra-chain position, and the position of adjacent mirror data units to obtain interleaved chain data, which specifically includes:
[0087] The initial chain data includes multiple chain sequences and the chain number and intra-chain position information of each mirror data unit in the corresponding chain sequence. The system processes each chain sequence one by one, determining the breakpoint insertion position according to a preset breakpoint generation strategy. This breakpoint generation strategy can be based on a fixed interval rule, i.e., setting a breakpoint position at a preset interval of mirror data units according to the intra-chain position; or it can be based on a chain length ratio rule or a local data change rule to determine the breakpoint position. After determining the breakpoint position, the system inserts a breakpoint identifier at the corresponding intra-chain position point, marking that position as a node in the chain structure that can be used for subsequent structural operations. The system binds each breakpoint identifier to the chain number of its chain sequence to clarify the chain to which the breakpoint belongs; records the intra-chain position of the breakpoint in the chain sequence to determine the specific location of the breakpoint in the chain; and reads the intra-chain position information of the mirror data units adjacent to the breakpoint, establishing a correspondence between these adjacent positions and the breakpoint identifier to describe the local structural interval where the breakpoint is located. The system updates or recalibrates the position order of the original mirrored data units within the chain. After completing the writing of breakpoints and the establishment of association relationships in all chain sequences, the system uniformly summarizes the data structure containing breakpoint identifiers, chain numbers, position order within the chain, and position order information of adjacent data units to form interleaved chain data.
[0088] In a preferred embodiment of the present invention, based on the interleaved chain data, data units from each data chain are introduced at each breakpoint, and a multi-chain interleaving network is constructed according to the intervals of the breakpoints to obtain interleaved structure data, including:
[0089] Based on the interleaved chain data, extract the breakpoint identifiers, chain numbers, and intra-chain positions in each data chain, calculate the interval positions between adjacent breakpoint identifiers, and obtain the breakpoint interval data.
[0090] Based on the breakpoint interval data, select the mirror data unit corresponding to the interval position according to the chain number, establish the position mapping relationship between the selected mirror data unit and the breakpoint identifier, and obtain the introduced unit data;
[0091] Based on the introduced unit data, the selected mirror data unit is written to the position of the breakpoint identifier, and the writing position is adjusted forward and backward according to the breakpoint interval data to obtain the embedded chain data.
[0092] Based on the embedded chain data, the cross-chain connection relationships formed by the breakpoint embedding in each data chain are recorded, and the cross-chain connection relationships are arranged and combined into a multi-chain intertwined relationship structure to obtain the intertwined structure data.
[0093] In this embodiment of the invention, based on the interleaved chain data, breakpoint identifiers, chain numbers, and intra-chain positions in each data chain are extracted. The interval position between adjacent breakpoint identifiers is calculated to obtain breakpoint interval data. Breakpoints are transformed from discrete position identifiers into structural information units, providing a basis for subsequent quantification. Based on the breakpoint interval data, mirror data units corresponding to the interval position are selected according to the chain number. A positional mapping relationship between the selected mirror data units and breakpoint identifiers is established to obtain introduced unit data. This allows mirror data units that originally existed independently in different chains to form a matching relationship through interval mapping. Based on the introduced unit data, the selected mirror data units are written to the position of the breakpoint identifier, and the writing position is adjusted forward and backward according to the breakpoint interval data to obtain embedded chain data. This allows data units that were originally distributed in different chains to form an intersecting relationship in the same chain structure, ensuring that the embedded position maintains the same distance relationship as the original chain. Based on the embedded chain data, the cross-chain connection relationships formed by breakpoint embedding in each data chain are recorded, and the cross-chain connection relationships are arranged and combined into a multi-chain interleaved relationship structure to obtain interleaved structure data. This transforms the connection between chains from an implicit relationship into an explicit structure, providing a data foundation for subsequent operations.
[0094] Specifically, based on the introduced unit data, the selected mirror data unit is written to the position of the breakpoint identifier, and the writing position is adjusted forward and backward according to the breakpoint interval data to obtain the embedded chain data, which includes:
[0095] The introduced unit data includes at least a breakpoint identifier, the chain number where the breakpoint is located, the breakpoint's position within the chain, the chain number to which the introduced unit belongs, the introduced unit's position within the chain, and the corresponding breakpoint interval value. The system uses the breakpoint identifier as the entry point, processing each breakpoint one by one to locate its specific position in the target chain. Specifically, it determines the breakpoint's index position in the chain sequence based on the chain number and its position within the chain. The system extracts the mirror data unit corresponding to the breakpoint from the introduced unit data and uses this mirror data unit as the object to be written. The system writes the introduced unit to the breakpoint position using either insertion or replacement. When using insertion, the system inserts the introduced unit before or after the breakpoint identifier, causing a partial expansion of the original chain structure. When using replacement, the introduced unit replaces the logical position of the breakpoint identifier, while retaining the breakpoint as auxiliary marker information. The system adjusts the position of the introduced unit based on the breakpoint interval data. The system reads the interval value corresponding to the breakpoint and uses this interval value as a position offset parameter to move the written introduced unit forward or backward. For example, when the interval value is large, the system can move the introduced unit backward along the chain's position order; when the interval value is small, it will keep it near the breakpoint or make minor adjustments forward. The system uniformly updates each chain sequence, forming a new chain structure containing original mirror data units and cross-chain introduced units. Each chain contains both original data units from its own chain and data units introduced from other chains, and the specific positions of these introduced units in the chain correspond to the breakpoint interval data. The system summarizes all the updated chain sequences to obtain the embedded chain data.
[0096] Specifically, based on the embedded chain data, the cross-chain connection relationships formed by breakpoint embedding in each data chain are recorded, and these cross-chain connection relationships are arranged and combined into a multi-chain intertwined relationship structure to obtain the intertwined structure data, which specifically includes:
[0097] The system reads the embedded chain data and traverses each chain sequence, identifying all cross-chain data units formed by breakpoint embedding operations. The system determines whether a data unit is a cross-chain introduction unit by comparing its source chain number with its current chain number. If they don't match, the data unit is identified as a data node generated by cross-chain embedding. For each identified cross-chain data unit, the system records its source chain number, target chain number (i.e., the current chain), its intra-chain position within the target chain, and its corresponding breakpoint identifier, forming a cross-chain connection relationship record. After scanning all chain sequences, the system obtains multiple sets of cross-chain connection relationships. This set is then structured and organized. Specifically, connections can be grouped according to the target chain number, representing all cross-chain introduction relationships within the same chain; grouped according to the source chain number to reflect the distribution of data units output from one chain to other chains; or sorted according to intra-chain position, giving the cross-chain relationships a clear sequential expression within the chain structure. The system represents all chain sequences and the sorted cross-chain connections as a unified multi-chain intertwined structure, and stores this structure data as intertwined structure data.
[0098] In a preferred embodiment of the present invention, based on the interleaved structure data, the frequency of occurrence of the same data unit in different positions and the cross-chain distribution range in each data chain are counted to identify the concentration of data units and obtain the structural aggregation quantity, including:
[0099] Based on the interleaved structure data, the number of times the same data unit appears in all data chains is counted to identify the density of the data unit in the multi-chain interleaved relationship structure and obtain the occurrence intensity term.
[0100] Based on the interleaved structure data, the span value between data chain numbers containing the same data unit is calculated, the cross-chain convergence degree of the same data unit is identified, and the distributed convergence term is obtained;
[0101] Based on the interleaved structure data, the difference between the maximum and minimum position of the data unit in the single data chain is calculated to identify the degree of positional clustering of the data unit in the single chain and obtain the clustering term in the chain.
[0102] Based on the interleaved structure data, calculate the position center value of the same data unit in each data chain, identify the degree of position consistency of the data unit between different data chains, and obtain the position consistency item;
[0103] By fusing the intensity term, distribution convergence term, intra-chain aggregation term, and positional consistency term, the concentration of data units in each data chain is identified, and the structural aggregation quantity is obtained.
[0104] In this embodiment of the invention, based on the interleaving structure data, the frequency of occurrence of the same data unit in all data chains is counted to identify the density of occurrence of the data unit in the multi-chain interleaving relationship structure, thus obtaining an occurrence intensity term to describe the density of occurrence of the data unit; based on the interleaving structure data, the span value between data chain numbers containing the same data unit is calculated to identify the cross-chain convergence degree of the same data unit, thus obtaining a distribution convergence term to reflect the distribution width of the data unit in the chain dimension, providing cross-chain distribution characteristics for subsequent processing; based on the interleaving structure data, the difference between the maximum and minimum position of the data unit in a single data chain is calculated. The algorithm identifies the degree of clustering of data units within a single chain, obtaining an intra-chain clustering term that describes the distribution concentration of data units within the chain, providing intra-chain structural features for subsequent analysis. Based on the interwoven structure data, it calculates the position center value of the same data unit in each data chain, identifying the degree of positional consistency of data units across different data chains, obtaining a positional consistency term that reflects the degree of positional coordination of data units in a multi-chain structure. Finally, it fuses the occurrence intensity term, distribution convergence term, intra-chain clustering term, and positional consistency term to identify the concentration of data units in each data chain, obtaining a structural aggregation quantity, and uniformly quantifying the overall concentration of data units.
[0105] The formula for calculating the amount of structural polymerization is as follows: ,
[0106] in, , , , ;
[0107] in, The occurrence intensity term represents the total occurrence intensity of the same data unit; the more times it appears, the larger the term becomes. This is the distribution convergence term, which represents the degree of clustering of the serial numbers of the same data unit across the chain. When the data unit is distributed across multiple chains but the chain numbers are consecutive, this term is larger; when the data unit is scattered and jumps between chains, this term is smaller. This is an intra-chain clustering term, representing the degree of clustering of the same data unit within each chain. The shorter the intra-chain expansion, the more concentrated the data is, and the larger this term is. The positional consistency term represents the degree of positional consistency of the same data unit in different data chains. The closer the normalized center position order is in different chains, the larger this term is.
[0108] in, The amount of polymer in the structure. A set of categories for the same data unit. For any of the same data unit categories, For containing data units The data chain set For the index of the data chain, For the first The total rank length within a data chain. For data units In the The number of occurrences in a data chain For data units Total number of occurrences in the entire data chain The maximum total number of occurrences across all data unit categories. For containing data units The number of data links, and Each contains a data unit The maximum and minimum chain numbers, Expand the width for the cross-chain number of the data unit. For data units In the The first in the data chain The order of occurrence and Data units In the The maximum and minimum bit order in a data chain. For data units In the The in-chain unwrap length in a data chain For data units In the The position center value in a data chain.
[0109] In a preferred embodiment of the present invention, based on the interleaved structure data, each data chain is divided into several segments according to the variation range of the data units, and the data units within the segments are folded to obtain folded segment data, including:
[0110] Based on the interleaved structure data, the intra-chain positional difference of adjacent data units in each data chain is calculated bit by bit, and the intra-chain positional difference is associated with the change of the identifier to obtain the change record data;
[0111] Based on the change record data, according to the change of identifiers between continuous data units, data units with the same change trend are divided into the same segment to obtain segmentation data;
[0112] Based on the segmented data, the data units within the same segment are merged and replaced with the original data units according to their arrangement order. At the same time, the position range of the merged data units is recorded to obtain the folded unit data.
[0113] Based on the folded unit data, each folded unit is rearranged according to its data chain and segment identifier, establishing a correspondence between the folded unit and the original data chain to obtain the folded segment data.
[0114] In this embodiment of the invention, based on the interleaved structure data, the intra-chain positional difference of adjacent data units in each data chain is calculated bit by bit, and the intra-chain positional difference is associated with the identifier change to obtain change record data. This transforms local changes in the chained data into a change record sequence, providing basic data for subsequent partitioning. Based on the change record data, according to the identifier change between consecutive data units, data units with the same change trend are divided into the same segment to obtain segmentation data. Data units are divided into multiple segments according to their change trends, transforming the chained data structure from point-by-point representation to segment representation. Based on the segmented data, data units within the same segment are merged and replaced with the original data units according to their arrangement order. At the same time, the position range of the merged data units is recorded to obtain folded unit data. The data representation in the chain structure is transformed from fine-grained units to segment-level units, reducing the number of data units and retaining segment boundary information. Based on the folded unit data, each folded unit is rearranged according to its data chain and segment identifier, establishing the correspondence between the folded unit and the original data chain to obtain folded segment data. This gives the segment structure a chain-like sequential relationship, achieving data compression and maintaining data traceability.
[0115] Specifically, based on the interleaved structure data, the intra-chain positional difference between adjacent data units in each data chain is calculated bit by bit, and the intra-chain positional difference is correlated with the identifier change to obtain change record data, which specifically includes:
[0116] The interwoven data structure includes at least multiple data chains, the order of data units within each data chain, the intra-chain positional information of each data unit, and identification information used to distinguish different data unit attributes. The system processes each data chain as a single data chain, scanning each chain bit-by-bit according to its intra-chain positional order from smallest to largest. During scanning, starting from the first valid data unit in the current chain, it sequentially selects adjacent data units to form consecutive pairs of adjacent data units. For each pair of adjacent data units, it reads the intra-chain positional value of the preceding and following data units, and calculates the intra-chain positional difference by subtracting the two. This positional difference characterizes the distance relationship or arrangement span of adjacent data units in the current chain. The system reads the corresponding identification information for each pair of adjacent data units and compares the two sets of identification information. The identification information can be a data unit identifier, category identifier, source identifier, structure identifier, or other fields used to characterize the attributes of the data units. During the comparison process, the system determines whether the identifiers of two consecutive data units are the same, or whether there is a change relationship between the identifiers under preset rules, such as changing from the same type of state to a different type of state, from the same source to a different source, or from the same structural category to another structural category. The system generates identifier change description information based on the comparison results. This identifier change description information can be represented by change markers, change type codes, change status values, or other recordable forms. The system associates the positional difference within the chain with the corresponding identifier change description information, forming a change record item for that adjacent data unit pair. The above operation is repeated along the current data chain until the positional difference calculation and identifier change association for all adjacent data unit pairs in the chain are completed. The same process is then performed on the next data chain until all data chains in the interleaved structure data have been scanned and recorded. All change record items formed in each data chain are then uniformly summarized to form change record data.
[0117] Specifically, based on the change record data, according to the changes in the identifiers between consecutive data units, data units with the same change trend are divided into the same segment, resulting in segmentation data, which specifically includes:
[0118] The system first uses the change record sequence corresponding to a single data chain as the analysis object for segment identification. Since the change record data of each data chain is already arranged according to the positional order within the chain, the system can start from the beginning of the chain and sequentially judge the change trend between adjacent data units along the change record sequence. The system selects the first change record item in the current chain as the initial reference record and extracts the identifier change description information from it. This identifier change description information is used as the initial change feature of the current segment. The system continues to read the next change record item in the chain and compares the identifier change description information corresponding to the change record item with the initial change feature of the current segment. When the system determines that the subsequent change record item is consistent with the initial change feature of the current segment, it means that the corresponding continuous data units maintain a continuous and consistent change pattern. At this time, the system continues to classify the data unit corresponding to the change record item into the current segment and synchronously updates the end position of the current segment. Consistency can be manifested as the identifier remaining unchanged, the identifier change type being the same, or multiple continuous change record items falling into the same preset change category. Conversely, when the system detects that the identifier change description information of a subsequent change record is inconsistent with the change characteristics of the current segment, it indicates that the change trend of the data unit in the chain has switched at that position. At this time, the system terminates the current segment and takes the data unit corresponding to the inconsistent position as the starting position of the new segment, and re-establishes the initial change characteristics of the segment.
[0119] The system assigns each segment a segment number, its chain number, the starting data unit's position within the chain, the ending data unit's position within the chain, the number of data units contained in the segment, and a segment change trend identifier. For segments at the beginning of a chain, the system uses the first data unit as the starting boundary; for segments at the end of a chain, after traversing to the last change record, the system includes its corresponding last data unit into the current segment to ensure that all data units in the entire chain are assigned to a specific segment. After traversing a data chain, the system obtains multiple consecutive segments with clearly defined boundaries. It processes the remaining data chains in the same way, ultimately summarizing the segment division results from all chains to form segment division data. Data units within each segment maintain a consistent change pattern, while boundaries between different segments are formed by switching change trends.
[0120] In a preferred embodiment of the present invention, based on the folded segment data, the source data chain of the data unit is traced step by step from the end of each segment, the number of path forks during the tracing process is recorded, the degree of data change and offset is identified, and the sequence offset is obtained, including:
[0121] Based on the folded segment data, the number of source forks in each folded segment is counted, the degree of diffusion of the folded segment in the source tracing process is identified, and the fork extension term is obtained;
[0122] Based on the folded segment data, calculate the intra-chain positional change of each folded segment, identify the degree of positional jump of the folded segment during the backtracking process, and obtain the positional jump term;
[0123] Based on the folded segment data, the activity level of cross-chain changes in the folded segment during the backtracking process is identified to obtain the inter-chain switching item;
[0124] By fusing the fork extension term, the positional shift term, and the inter-chain switching term, the overall change and offset of the folded segment during the backtracking process are identified, and the sequence offset is obtained. In this embodiment of the invention, based on the folded segment data, the number of source forks in each folded segment is counted, the diffusion degree of the folded segment in the source tracing process is identified, and a fork extension term is obtained, which quantifies the multi-path source relationship originally implicit in the data reorganization process, forming a description of the diffusion degree of data source; based on the folded segment data, the intra-chain positional change of each folded segment is calculated, the positional jump degree of the folded segment in the backtracking process is identified, and a positional jump term is obtained, which transforms the positional migration of data in the chain structure from discrete positional differences into numerical indicators, forming an expression of the positional jump degree; based on the folded segment data, the activity level of cross-chain changes of the folded segment in the backtracking process is identified, and an inter-chain switching term is obtained, which quantifies the migration behavior of data in the multi-chain structure, forming an expression of the activity level of cross-chain changes; the fork extension term, positional jump term, and inter-chain switching term are fused to identify the overall change offset degree of the folded segment in the backtracking process, and a sequence offset is obtained, which quantifies the overall change offset degree of the data.
[0125] The formula for calculating the sequence offset is as follows: ,
[0126] in, , , ;
[0127] in, The bifurcation extension term incorporates the number of source bifurcations in each backtracking level and the level depth into the calculation. This term is larger as the source becomes more dispersed when tracing back to deeper levels, representing the instability of the source path in the folded segment. For the positional shift term, calculate the intra-chain positional difference between adjacent backtracking levels and normalize it using the positional range length of this segment. The larger the term is, the more obvious the positional shift during the backtracking process, which characterizes the degree of positional shift of the data unit in the backtracking chain. For the inter-chain switching term, the change in the source chain number between adjacent backtracking levels is calculated and normalized by the maximum number difference. This term is larger when the cross-chain switching is more frequent and obvious during the backtracking process, representing the degree of cross-chain disturbance in the data source path.
[0128] in, This is the sequence offset. For a set of folded segments; For the index of the folded section; For the first The number of backtracking levels for each folded segment; For the first The number of source forks of each folded segment at the i-th backtracking level; The maximum number of source forks across all backtracking levels; For the first The folding section in the first The intra-chain positional order corresponding to each backtracking level; For the first The length of the positional range of each folded segment; For the first The folding section in the first The source chain number corresponding to each backtracking level; The maximum number difference between all source chain numbers; The minimum value is taken as the preset parameter.
[0129] In a preferred embodiment of the present invention, by mapping structural aggregation and sequence offset to aggregation level intervals and offset level intervals, combining different level intervals and establishing rule encoding, rule data is obtained, including:
[0130] By dividing the structural aggregation quantity into multiple aggregation level intervals according to preset aggregation values, and assigning an aggregation interval identifier to each aggregation level interval, aggregation interval data is obtained.
[0131] By dividing the sequence offset into multiple offset level intervals according to preset offset values and assigning an offset interval identifier to each offset level interval, offset interval data is obtained.
[0132] By merging each aggregation level interval with each offset level interval into a combined interval, and establishing a combined identifier for each combined interval, the interval combined data is obtained;
[0133] Based on the interval combination data, the combination identifier is converted into a rule code according to the preset encoding rules, and a mapping relationship between the rule code and the combination interval is established to obtain the rule data.
[0134] In this embodiment of the invention, by dividing the structural aggregation quantity into multiple aggregation level intervals according to a preset aggregation value and assigning an aggregation interval identifier to each aggregation level interval, aggregation interval data is obtained. The numerical value of rule matching is transformed into a category identifier in a finite set, providing input for subsequent combination encoding. By dividing the sequence offset into multiple offset level intervals according to a preset offset value and assigning an offset interval identifier to each offset level interval, offset interval data is obtained, so that the path change feature is expressed in a finite category form, forming a data representation structure with the aggregation interval. By merging each aggregation level interval and each offset level interval into a combined interval and establishing a combined identifier for each combined interval, interval combined data is obtained, expanding the state of the data object from a single-dimensional classification to a multi-dimensional joint classification, forming a data category system with discriminative power. Based on the interval combined data, the combined identifier is converted into a rule code according to a preset encoding rule, and a mapping relationship between the rule code and the combined interval is established to obtain rule data, realizing the structured expression of the data state and providing input for subsequent content matching.
[0135] Specifically, by dividing the structural aggregation quantity into multiple aggregation level intervals according to preset aggregation values and assigning an aggregation interval identifier to each aggregation level interval, aggregation interval data is obtained, which includes:
[0136] The structural aggregation quantity corresponds to each data unit or folded segment, and is numerically characterized by its distribution concentration in the multi-chain interwoven structure. The system determines the overall numerical range of the structural aggregation quantity, for example, by traversing all structural aggregation quantities to obtain the minimum, maximum, and intermediate distribution. The system then segments this numerical range according to preset aggregation value division rules. These rules can employ a fixed threshold method, i.e., pre-setting multiple boundary values to divide the numerical range into several continuous intervals; or they can employ a dynamic division method based on data distribution, such as dividing interval boundaries according to the statistical distribution density or quantile of the structural aggregation quantity. The system assigns a unique aggregation interval identifier to each aggregation level interval. This identifier can be a sequential number, character encoding, or other form of marker. The system compares the structural aggregation quantity corresponding to each data object with the aforementioned interval boundaries to determine the specific interval range it falls into, and writes the corresponding aggregation interval identifier into the attribute record of the data object. The system synchronously records the mapping relationship between the original aggregation quantity value of the data object and the interval identifier for subsequent backtracking or analysis. The system then organizes the data set containing the correspondence between data objects, structural aggregation quantities, and aggregation interval identifiers into unified aggregate interval data.
[0137] Specifically, by dividing the sequence offset into multiple offset level intervals according to preset offset values and assigning an offset interval identifier to each offset level interval, offset interval data is obtained, which includes:
[0138] The sequence offset is used to characterize the degree of path change of each folded segment during the source tracing process. The system traverses all sequence offsets, statistically analyzes their numerical distribution range, and determines the basis for dividing the offset intervals based on this range. The sequence offsets are segmented according to a preset offset value division rule. Specifically, an equal-interval division method can be used to evenly divide the overall range into several intervals; alternatively, non-uniform intervals can be set based on the density of the sequence offset distribution, for example, setting more subdivided interval ranges in intervals where numerical changes are more concentrated. The system assigns a corresponding offset interval identifier to each offset level interval. This identifier can also be represented by a number or code. For each data object, its corresponding sequence offset is matched with the boundaries of each interval to determine its offset level interval, and the offset interval identifier is written into the attribute record of the data object. Simultaneously, the correspondence between the original values of the sequence offsets and the interval identifiers is recorded to form a traceable data mapping structure. The system then uniformly summarizes the data containing the relationships between data objects, sequence offsets, and offset interval identifiers to obtain the offset interval data.
[0139] Specifically, based on the interval combination data, the combination identifier is converted into a rule code according to a preset encoding rule, and a mapping relationship is established between the rule code and the combination interval to obtain the rule data, which includes:
[0140] In the interval combination data, each data object corresponds to a combined interval, which is composed of its aggregation interval identifier and offset interval identifier. For each combined interval, the system extracts its combination identifier and performs conversion processing according to preset encoding rules. The encoding rules can be structured encoding rules, such as mapping the aggregation interval identifier to the first part of the code and the offset interval identifier to the second part of the code, concatenating them in a preset order to generate a rule code; alternatively, a lookup table method can be used, that is, by pre-establishing a correspondence table between combined intervals and rule codes, directly mapping the combination identifier to the corresponding code value. During the encoding generation process, the system standardizes the encoding, such as unifying the code length, unifying the code format, or adding check bits, to ensure the consistency and recognizability of the rule codes in the system. The system establishes a mapping relationship between each generated rule code and its corresponding combined interval and records this mapping relationship in a rule mapping table. This mapping table contains at least fields such as rule code, aggregation interval identifier, offset interval identifier, and combined interval identifier, used to describe the correspondence between the code and the original interval combination. The system summarizes all rule codes and their mapping relationships to form rule data. This rule data serves as direct input for subsequent content matching, enabling multidimensional statistical results to be transformed into directly accessible structured data through encoding.
[0141] Embodiments of the present invention also provide a digital content adaptive generation system based on scene semantic understanding, the system comprising:
[0142] The mirroring module is used to obtain the input dataset, copy each data unit in the input dataset to generate at least two mirror data units, and assign different sequence bias labels to each mirror data unit to obtain the mirror dataset;
[0143] The interleaving module is used to alternately insert mirror data units with different sequence bias identifiers into multiple data chains based on the mirror dataset, and set breakpoint positions in each data chain to obtain interleaved chain data;
[0144] The interleaving module is used to introduce data units from each data chain at each breakpoint position based on the interleaved chain data, and construct a multi-chain interleaving network according to the interval of the breakpoint positions to obtain interleaved structure data.
[0145] The aggregation module is used to count the number of times the same data unit appears in different positions in each data chain and the cross-chain distribution range based on the interwoven structure data, identify the concentration of data units, and obtain the structure aggregation quantity.
[0146] The folding module is used to divide each data chain into several segments according to the variation range of data units based on the interleaved structure data, and fold the data units within the segments to obtain folded segment data;
[0147] The offset module is used to trace the source data chain of the data unit step by step from the end of each segment based on the folded segment data, record the number of path forks during the tracing process, identify the degree of data change and offset, and obtain the sequence offset.
[0148] The rules module is used to map structural aggregation quantities and sequence offsets to aggregation level intervals and offset level intervals, combine different level intervals and establish rule codes to obtain rule data;
[0149] The matching module is used to split the preset content resources into multiple content units according to the rule data, and to encode the matching content units for different rules to obtain the target content data.
[0150] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.
[0151] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0152] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0153] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A digital content adaptive generation method based on scene semantic understanding, characterized in that, The method includes: Obtain the input dataset, copy each data unit in the input dataset to generate at least two mirror data units, and assign different sequence bias labels to each mirror data unit to obtain the mirror dataset; Based on the mirror dataset, mirror data units with different sequence bias identifiers are alternately inserted to construct multiple data chains, and breakpoint positions are set in each data chain to obtain interleaved chain data; Based on the interleaved chain data, data units from each data chain are introduced at each breakpoint, and a multi-chain interleaved network is constructed according to the interval of the breakpoint to obtain the interleaved structure data. Based on the interleaved structure data, the frequency of occurrence of the same data unit in different positions and the cross-chain distribution range in each data chain are counted to identify the concentration of data units and obtain the structural aggregation quantity, including: Based on the interleaved structure data, the number of times the same data unit appears in all data chains is counted to identify the density of the data unit in the multi-chain interleaved relationship structure and obtain the occurrence intensity term. Based on the interleaved structure data, the span value between data chain numbers containing the same data unit is calculated, the cross-chain convergence degree of the same data unit is identified, and the distributed convergence term is obtained; Based on the interleaved structure data, the difference between the maximum and minimum position of the data unit in the single data chain is calculated to identify the degree of positional clustering of the data unit in the single chain and obtain the clustering term in the chain. Based on the interleaved structure data, calculate the position center value of the same data unit in each data chain, identify the degree of position consistency of the data unit between different data chains, and obtain the position consistency item; By fusing the intensity term, distribution convergence term, intra-chain clustering term, and positional consistency term, the concentration of data units in each data chain is identified, and the structural aggregation quantity is obtained. Based on the interleaved structure data, each data chain is divided into several segments according to the variation range of the data units, and the data units within the segments are folded to obtain folded segment data; Based on the folded segment data, the source data chain of the data unit is traced step by step from the end of each segment. The number of path forks during the tracing process is recorded, the degree of data change and offset is identified, and the sequence offset is obtained, including: Based on the folded segment data, the number of source forks in each folded segment is counted, the degree of diffusion of the folded segment in the source tracing process is identified, and the fork extension term is obtained; Based on the folded segment data, calculate the intra-chain positional change of each folded segment, identify the degree of positional jump of the folded segment during the backtracking process, and obtain the positional jump term; Based on the folded segment data, the activity level of cross-chain changes in the folded segment during the backtracking process is identified to obtain the inter-chain switching item; By fusing the fork extension term, the positional shift term, and the inter-chain switching term, the overall change and offset of the folded segment during the backtracking process is identified, and the sequence offset is obtained. By mapping structural aggregation and sequence offset to aggregation level intervals and offset level intervals, and combining different level intervals to establish rule encoding, rule data is obtained. Based on the rule data, the preset content resources are divided into multiple content units, and the content units are coded and matched for different rules to obtain the target content data.
2. The adaptive digital content generation method based on scene semantic understanding according to claim 1, characterized in that, Based on the mirror dataset, mirror data units with different sequence bias identifiers are alternately inserted to construct multiple data chains, and breakpoints are set in each data chain to obtain interleaved chain data, including: Based on the mirror dataset, each mirror data unit is divided into multiple bias unit groups according to the sequence bias identifier, and the order of each mirror data unit in the bias unit group is recorded to obtain the bias grouped data. Based on the bias grouped data, the order of the mirrored data units within the same bias unit group is maintained, and the mirrored data units are written sequentially into the sorting queue to obtain the alternating queue data. Based on the alternating queue data, the mirror data units in the permutation queue are assigned to different chain sequences, and the chain number and position within the chain of each mirror data unit are recorded to obtain the initial chain data; Based on the initial chain data, breakpoint identifiers are written into each chain sequence, and the breakpoint identifiers are associated with the chain number, the position within the chain, and the position of adjacent mirror data units to obtain interleaved chain data.
3. The adaptive digital content generation method based on scene semantic understanding according to claim 2, characterized in that, Based on the interleaved chain data, data units from each data chain are introduced at each breakpoint, and a multi-chain interleaved network is constructed according to the intervals between the breakpoints to obtain the interleaved structure data, including: Based on the interleaved chain data, extract the breakpoint identifiers, chain numbers, and intra-chain positions in each data chain, calculate the interval positions between adjacent breakpoint identifiers, and obtain the breakpoint interval data. Based on the breakpoint interval data, select the mirror data unit corresponding to the interval position according to the chain number, establish the position mapping relationship between the selected mirror data unit and the breakpoint identifier, and obtain the introduced unit data. Based on the introduced unit data, the selected mirror data unit is written to the position of the breakpoint identifier, and the writing position is adjusted forward and backward according to the breakpoint interval data to obtain the embedded chain data. Based on the embedded chain data, the cross-chain connection relationships formed by the breakpoint embedding in each data chain are recorded, and the cross-chain connection relationships are arranged and combined into a multi-chain intertwined relationship structure to obtain the intertwined structure data.
4. The adaptive digital content generation method based on scene semantic understanding according to claim 3, characterized in that, Based on the interleaved structure data, each data chain is divided into several segments according to the variation range of the data units, and the data units within each segment are folded to obtain folded segment data, including: Based on the interleaved structure data, the intra-chain positional difference of adjacent data units in each data chain is calculated bit by bit, and the intra-chain positional difference is associated with the change of the identifier to obtain the change record data; Based on the change record data, according to the change of identifiers between continuous data units, data units with the same change trend are divided into the same segment to obtain segmentation data; Based on the segmented data, the data units within the same segment are merged and replaced with the original data units according to their arrangement order. At the same time, the position range of the merged data units is recorded to obtain the folded unit data. Based on the folded unit data, each folded unit is rearranged according to its data chain and segment identifier, establishing a correspondence between the folded unit and the original data chain to obtain the folded segment data.
5. The adaptive digital content generation method based on scene semantic understanding according to claim 4, characterized in that, By mapping structural aggregation quantities and sequence offsets to aggregation level intervals and offset level intervals, and combining different level intervals to establish rule encoding, rule data is obtained, including: By dividing the structural aggregation quantity into multiple aggregation level intervals according to preset aggregation values, and assigning an aggregation interval identifier to each aggregation level interval, aggregation interval data is obtained. By dividing the sequence offset into multiple offset level intervals according to preset offset values and assigning an offset interval identifier to each offset level interval, offset interval data is obtained. By merging each aggregation level interval with each offset level interval into a combined interval, and establishing a combined identifier for each combined interval, the interval combined data is obtained; Based on the interval combination data, the combination identifier is converted into a rule code according to the preset encoding rules, and a mapping relationship between the rule code and the combination interval is established to obtain the rule data.
6. A digital content adaptive generation system based on scene semantic understanding, characterized in that, The system is used to perform the method as described in any one of claims 1 to 5, the system comprising: The mirroring module is used to obtain the input dataset, copy each data unit in the input dataset to generate at least two mirror data units, and assign different sequence bias labels to each mirror data unit to obtain the mirror dataset; The interleaving module is used to alternately insert mirror data units with different sequence bias identifiers into multiple data chains based on the mirror dataset, and set breakpoint positions in each data chain to obtain interleaved chain data; The interleaving module is used to introduce data units from each data chain at each breakpoint position based on the interleaved chain data, and construct a multi-chain interleaving network according to the interval of the breakpoint positions to obtain interleaved structure data. The aggregation module is used to count the number of times the same data unit appears in different positions in each data chain and the cross-chain distribution range based on the interwoven structure data, identify the degree of concentration of data units, and obtain the structure aggregation quantity. The folding module is used to divide each data chain into several segments according to the variation range of data units based on the interleaved structure data, and fold the data units within the segments to obtain folded segment data; The offset module is used to trace the source data chain of the data unit step by step from the end of each segment based on the folded segment data, record the number of path forks during the tracing process, identify the degree of data change and offset, and obtain the sequence offset. The rules module is used to map structural aggregation quantities and sequence offsets to aggregation level intervals and offset level intervals, combine different level intervals and establish rule codes to obtain rule data; The matching module is used to split the preset content resources into multiple content units according to the rule data, and to encode the matching content units for different rules to obtain the target content data.
7. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Abstract generation and analysis method and device of target report, equipment and medium
CN121188773A
Semantic perception black box large language model training data auditing method and system
CN121256815A