An online editing method and system for multimedia documents
By constructing the initial structure diagram of the graphics and text and calculating the difference between the offset angle and the path, and collecting the time series of editing behavior, the problem of image and text misalignment in online editing of multimedia documents is solved, and the stability and response efficiency of the document structure are improved.
Patent Information
- Application Number
- CN202510526981.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The prior art lacks the ability to dynamically parse relative coordinates and offset paths between pictures and texts in online editing of multimedia documents, resulting in misalignment of images and text and reconstruction difficulties, making it difficult to restore operation timing and trigger chains, causing confusion in document structure and affecting stability and interaction coherence.
By obtaining the image boundary coordinates and text starting position, building the initial structure diagram of the graphic and text, calculating the offset angle and path difference, generating a path offset identification set, collecting the time series and label features of the editing behavior, calculating the task synchronization rhythm, reconstructing the logical order of text blocks, and realizing the rhythm control of content loading.
It improves the stability and response efficiency of multimedia documents in complex editing scenarios, maintains the consistency and precise positioning of structure re-arrangement, and enhances the stability of multi-user collaboration and cross-terminal editing.
Smart Images

Figure CN120068817B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of document processing, and particularly to an online editing method and system for multimedia documents. Background Art
[0002] The technical field of document processing includes the whole process of document creation, editing, management, retrieval, and storage, covering comprehensive processing methods for multimedia content such as text, images, audio, and video. The core content of this technical field lies in realizing the unified access, structured processing, and visual operation of multi-source heterogeneous data, and supporting key links such as collaborative work, permission control, and version management. With the development of information technology, document processing technology has gradually shifted from local processing to online and intelligent processing, and is widely used in fields such as office automation, content management systems, and digital archive management. It has the ability to efficiently process large-scale document data and support remote real-time collaboration, reflecting a strong dependence on the integration of data organization and content services.
[0003] Among them, the online editing method for multimedia documents refers to a technical method for composite documents containing elements such as text, pictures, audio, and video, which provides an editing interface through a network platform to support remote real-time collaborative editing. This method usually covers multiple technical matters such as format parsing, unified rendering, logical structure recognition, cross-terminal compatibility processing, and cloud data synchronization of multimedia content. Specific methods include marking and parsing multiple types of content based on a predefined document structure model, using a web editing kernel to perform online loading and editing control of multimedia elements, and completing real-time synchronization of document status and transmission and storage of modified data through the communication protocol between the server and the client. Such methods generally take the modeling of the document content structure as the basis, and combine real-time network communication mechanisms and browser visualization interfaces to achieve concurrent access and editing by multiple users.
[0004] In the prior art, when dealing with the combination relationship between graphics and text, it relies on a static structure model for position determination, lacking the ability to dynamically analyze the relative coordinates and offset paths between graphics and text, resulting in misalignment and difficult reconstruction of images and text in the document. In the face of high-frequency and multi-dimensional operation inputs from users, it fails to effectively model the operation sequence and trigger chain, making it difficult to restore the causal relationship before and after in the operation chain, resulting in problems such as chaotic structure update logic and misordered edited content. In the scenario of task concurrency and audio-video mixed content, the current mechanism lacks quantitative control over the synchronization degree of content loading and operation response, easily causing editing stagnation or task blocking. In the structural adjustment stage, existing methods mostly perform simple rearrangement based on the content presentation order, lacking fine calculation of logical paths and structural spacing, resulting in problems such as logical breakage and disordered combination relationships after the document structure is rearranged. These deficiencies are particularly prominent in application environments such as multi-user collaboration, cross-terminal editing, and complex graphic-text mixed layout, directly affecting the structural stability of document content and the interaction coherence during use. Summary of the Invention
[0005] The object of the present invention is to solve the disadvantages existing in the prior art, and a method and system for online editing of multimedia documents are proposed.
[0006] In order to achieve the above object, the present invention adopts the following technical solution: A method for online editing of multimedia documents, including the following steps:
[0007] S1: Obtain the image boundary coordinates and the starting position of the text, extract the difference in combination with the graphic-text alignment direction, generate a path sequence based on the offset of adjacent elements, match the combined tags and the editing attribute status, and construct an initial graphic-text structure diagram;
[0008] S2: Call the path of the initial structure diagram, compare the current graphic-text coordinates with the original path vector, calculate the offset angle and the path difference through the scaling ratio and the scrolling displacement, and generate a path offset identification set;
[0009] S3: Call the graphic-text combination index in the offset identification set, collect the timestamp, layer level and action tags during user editing, calculate the trigger interval between actions and the tag fitting degree, screen continuous operation nodes, extract the trigger order within the graphic-text combination, and generate a path offset trigger linked list;
[0010] S4: Call the audio-visual content types associated with the continuous operation nodes in the path offset trigger linked list, obtain the rendering delay and the response time difference, calculate the task synchronization deviation, determine whether it meets the channel multiplexing tolerance interval, and generate a reusable task index set;
[0011] S5: According to the reusable task index set, locate the corresponding elements in the editing interface, compare their combined distance with the original order, reconstruct the logical order of the text blocks, and generate the rearrangement result of the multimedia document structure.
[0012] As a further solution of the present invention, the initial graphic-text structure diagram includes image boundary coordinates, text character positions, path direction information, combined tag types, and editing attribute statuses. The path offset identification set includes offset angle values, path distance differences, scaling ratio parameters, and scrolling displacement amounts. The path offset trigger linked list includes operation time nodes, layer level identifiers, action type tags, and trigger order information. The reusable task index set includes audio-visual content types, rendering delay parameters, response time difference indicators, and channel multiplexing tolerance ranges. The rearrangement result of the multimedia document structure includes element combination distances, text block order relationships, associated logical paths, and interface positioning positions.
[0013] As a further solution of the present invention, the specific steps of S1 are:
[0014] S101: Obtain the pixel coordinates of the image boundary, match the starting position of the text characters, calculate the offset between the two, extract the coordinate differences in the horizontal and vertical directions, establish a coordinate matching relationship, and obtain the pixel-text matching offset;
[0015] S102: Based on the pixel-text matching offset, extract the offset directions of adjacent characters, calculate the displacement difference, filter the data that conforms to the text block alignment direction, generate a character connection path, and call the offset data for path sorting to obtain a character offset path sequence;
[0016] S103: Call the character offset path sequence, match the graphic-text combination label with the editing attribute status identifier, analyze the path change trend, filter the combination methods that conform to the structure, adjust the graphic-text association relationship, and obtain the initial graphic-text structure diagram.
[0017] As a further solution of the present invention, the specific steps of S2 are as follows:
[0018] S201: Call the graphic-text combination path in the initial graphic-text structure diagram, obtain the coordinate positions of the graphic-text elements in the current view, calculate the difference between the current view coordinates and the original path vector, establish a corresponding relationship between the two, and obtain the graphic-text coordinate comparison data;
[0019] S202: Based on the graphic-text coordinate comparison data, calculate the overall scaling and scrolling offset index, extract the offset angle and path distance difference between adjacent data points during the coordinate transformation process, analyze the path change trend, filter the change data in the path offset direction, and obtain the path offset calculation result;
[0020] S203: Call the path offset calculation result, analyze the structural changes of the combination path, identify the adjustment status of the graphic-text combination relationship, filter the feature points of the changed path, and extract the corresponding offset mode to obtain a path offset identification set.
[0021] As a further solution of the present invention, the specific formula for the overall scaling and scrolling offset index is as follows:
[0022] ;
[0023] Wherein, represents the overall scaling and scrolling offset index, represents the horizontal coordinate value of the th coordinate point in the current view, represents the horizontal coordinate value of the th coordinate point in the original view, represents the vertical coordinate value of the th coordinate point in the current view, represents the vertical coordinate value of the th coordinate point in the original view, Represents the path distance of the nth coordinate point in the original view, A small positive number to prevent the denominator from being zero, Represents the total number of coordinate data points, Represents the nth scrolling displacement value in the current view, Represents the nth scrolling displacement value in the original view, Represents the total number of scrolling displacement data points.
[0024] As a further solution of the present invention, the specific steps of S3 are as follows:
[0025] S301: Call the path offset identification set, extract the graphic-text combination indexes with path offsets, screen the corresponding graphic-text combination data, establish an index list, and obtain a graphic-text offset index set;
[0026] S302: Based on the graphic-text offset index set, collect the timestamps, layer levels, and action tags in the user's editing operations, calculate the average trigger time interval, screen the action sequences that conform to the path offset rule, analyze the fitting degree between the tags, identify the association relationship between the actions, and obtain action trigger matching data;
[0027] S303: Call the action trigger matching data, construct a continuous operation node link, extract the trigger order within the graphic-text combination, screen the key nodes with path offsets, and obtain a path offset trigger linked list.
[0028] As a further solution of the present invention, the specific formula for the average trigger time interval is:
[0029] ;
[0030] Wherein, Represents the average trigger time interval of the nth operation relative to the previous operation, Represents the timestamp of the nth operation in the current state, Represents the timestamp of the nth operation, Represents the adjustment coefficient of the layer level change, Represents the layer level of the nth layer in the current state, Represents the layer level of the mth layer, Represents the trigger path offset of the nth operation in the current state, Represents the trigger path offset of the th operation in the original state, represents the total number of path offset data points involved in the interface.
[0031] As a further solution of the present invention, the specific steps of S4 are as follows:
[0032] S401: Call the path offset trigger linked list, extract consecutive operation nodes, obtain the corresponding audio-visual content types, screen the content data with synchronizable characteristics, and obtain an audio-visual type matching set;
[0033] S402: Based on the audio-visual type matching set, calculate the rendering delay and the operation response time difference, extract the task loading timeliness data, calculate the synchronization difference between task loading and response, screen the task nodes that meet the synchronization characteristics, and obtain the task synchronization difference calculation result;
[0034] S403: Call the task synchronization difference calculation result, judge the compliance of the channel multiplexing tolerance interval, extract the task indexes that meet the multiplexing conditions, screen the reusable task data, and obtain a reusable task index set.
[0035] As a further solution of the present invention, the specific steps of S5 are as follows:
[0036] S501: Call the reusable task index set, extract the graphic and text combination information, identify the corresponding element positions in the editing interface, screen all the matched combined elements, and obtain the graphic and text combination positioning data;
[0037] S502: Based on the graphic and text combination positioning data, calculate the combination distance between the current element position and the original order, screen the text blocks that meet the logical distance requirements, extract the recombination sequence of the text blocks, and call the element order matched with the combination distance data to obtain the text combination reconstruction sequence;
[0038] S503: Call the text combination reconstruction sequence, re-establish the associated logical path, screen the adjusted graphic and text combination structure, extract the current document element sorting, and obtain the multimedia document structure rearrangement result.
[0039] A multimedia document online editing system, comprising:
[0040] The graphic and text structure analysis module obtains the pixel coordinates of the image boundary and the starting position of the corresponding text characters, calculates the coordinate difference in the alignment direction of the graphic and text blocks, generates a path sequence according to the offset direction of adjacent elements, matches the graphic and text combination tags and the editing attribute status identifiers, identifies the graphic and text combination relationship, and obtains the initial graphic and text structure diagram;
[0041] The view offset detection module calls the initial structure diagram of graphics and text, detects the coordinate positions of graphics and text elements in the view, compares them with the original path vector, calculates the scaling ratio, scrolling displacement, and offset angle, determines the change status of the combination relationship, and obtains a set of path offset identifiers;
[0042] The editing behavior analysis module calls the set of path offset identifiers, extracts the graphics and text combination index, collects the timestamps, layer levels, and action labels of the user's editing operations, calculates the trigger intervals and label fitting degrees of adjacent actions, constructs a continuous operation node link, and generates a path offset trigger linked list;
[0043] The task synchronization calculation module calls the path offset trigger linked list, analyzes the operations of audio and video content, calculates the rendering delay and the time difference between operation responses, compares the synchronization difference between the loading timeliness and the response rhythm, determines the channel multiplexing tolerance interval, and generates a set of reusable task indexes;
[0044] The document structure rearrangement module calls the set of reusable task indexes, extracts the positions of graphics and text elements in the editing interface, calculates the combination distance index from the original order, filters the text blocks with matching logical distances, reconstructs the graphics and text combination order and the associated logical path, and generates the result of multimedia document structure rearrangement.
[0045] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0046] In the present invention, a logical structure is constructed through the coordinate difference of graphics and text and the offset path, the change of the combination state is accurately judged by combining the angle and distance changes, the time series and label features in the editing behavior are collected, an operation link is established to improve the timing consistency of content recombination, the task synchronization rhythm is calculated through the delay and response difference, the rhythm control of content loading is realized, the combination order is reconstructed according to the logical distance, the coherence and accurate positioning after structure rearrangement are maintained, and the stability and response efficiency of the multimedia document in complex editing scenarios are enhanced. Description of the Drawings
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0048] Figure 1 It is a schematic flow chart of the steps of the present invention;
[0049] Figure 2 It is a system module diagram of the present invention. Detailed Embodiments
[0050] The technical solutions in the present invention will be described below with reference to the accompanying drawings.
[0051] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "example" in the present invention should not be construed as more preferred or more advantageous than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.
[0052] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same. "Of", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same.
[0053] In the embodiments of the present invention, sometimes subscripts such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meanings they express are the same.
[0054] To make the technical problems to be solved, technical solutions and advantages of the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0055] Please refer to Figure 1 , a method for online editing of multimedia documents, including the following steps:
[0056] S1: Obtain the pixel coordinates of the image boundary and the starting position of the corresponding text character, extract the coordinate difference in combination with the alignment direction of the graphic and text block, generate a path sequence based on the offset direction between adjacent elements, and call the graphic and text combination tag and the editing attribute status identifier for combined matching to obtain the initial graphic and text structure diagram;
[0057] S2: Call the graphic and text combination path in the initial graphic and text structure diagram, compare the coordinate positions of the graphic and text elements in the current view with the original path vector, calculate the offset angle and the path distance difference through the zoom ratio and the scroll displacement, and identify the change state of the combination relationship to obtain the path offset identifier set;
[0058] S3: Call the graphic and text combination index with path offset in the path offset identifier set, collect the timestamp, layer level, and action tag in the user's editing operation, calculate the trigger interval and the tag fitting degree between adjacent actions, form a continuous operation node link, extract the trigger order within the graphic and text combination, and generate a path offset trigger linked list;
[0059] S4: Invoke the audio - video content types associated with consecutive operation nodes in the path offset trigger linked list, obtain the rendering delay and the operation response time difference, calculate the synchronization difference between the current task's loading timeliness and response rhythm, and determine the compliance with the channel multiplexing tolerance interval to generate a reusable task index set;
[0060] S5: According to the graphic - text combination information in the reusable task index set, identify the corresponding element positions in the editing interface, compare the combination distance between the positions and the original order, select the text blocks corresponding to the logical distance to reconstruct the combination order, re - establish the associated logical path, and generate the multimedia document structure rearrangement result.
[0061] The initial graphic - text structure diagram includes image boundary coordinates, text character positions, path direction information, combination label types, and editing attribute statuses. The path offset identifier set includes offset angle values, path distance differences, scaling ratio parameters, and scrolling displacement amounts. The path offset trigger linked list includes operation time nodes, layer - level identifiers, action type labels, and trigger order information. The reusable task index set includes audio - video content types, rendering delay parameters, response time difference indicators, and channel multiplexing tolerance ranges. The multimedia document structure rearrangement result includes element combination distances, text block order relationships, associated logical paths, and interface positioning positions.
[0062] The specific steps of S1 are as follows:
[0063] S101: Obtain the pixel coordinates of the image boundary, match the starting positions of the text characters, calculate the offset between the two, extract the coordinate differences in the horizontal and vertical directions, establish a coordinate matching relationship, and obtain the pixel - text matching offset.
[0064] The input image needs to be preprocessed first, including grayscale conversion and binarization, in order to highlight the text and boundary information in the image. On this basis, by traversing the pixel points of the image, regions with significant brightness changes are screened as possible boundary regions, and the contour lines of the image are further extracted through edge detection methods. After obtaining the boundary pixel coordinates, they are stored as an ordered coordinate set. At the same time, the starting positions of text characters are located. Using optical character recognition (OCR) technology, the characters in the image are scanned and the upper left coordinate points of each character are output, and the starting positions of all characters are extracted. Subsequently, the horizontal and vertical deviations between the boundary pixel coordinates of the image and the starting positions of text characters are calculated. The horizontal deviation value represents the position difference of the characters in the horizontal direction, and the vertical deviation value represents the position difference of the characters in the vertical direction. By traversing the deviation values of all characters, the distribution range of the deviation values is statistically analyzed, and the values with a smaller change range are screened to determine the matching offset of the pixel text. For example, if the horizontal offset value of the character starting coordinates is stable between 30 and 35 pixels, and the vertical offset value is between 5 and 10 pixels, it can be determined that the pixel text offset of this image is mainly concentrated in these two ranges. If the offset values of some characters are far from this range, it is necessary to judge whether they are abnormal data and further adjust the matching rules. Finally, the matching relationship between pixel coordinates and text coordinates is established to determine the matching offset.
[0065] S102: Based on the matching offset of pixel text, extract the offset directions of adjacent characters, calculate the displacement difference, screen the data that conforms to the text block alignment direction, generate a character connection path, and call the offset data to perform path sorting to obtain a character offset path sequence;
[0066] Extract the offset directions of adjacent characters. First, arrange the characters in the order from left to right and from top to bottom, calculate the coordinate changes between every two adjacent characters, count the horizontal offset values and vertical offset values of the characters, and filter the data that conforms to the alignment direction of the text block. For example, if the text is arranged in standard left alignment, the horizontal offset values of all adjacent characters should be within a certain range, such as between 20 and 25 pixels, and the vertical offset value is close to zero, indicating that the text is regularly arranged in the horizontal direction. If the vertical offset value is large, it indicates that there may be line spacing, and the structure of the text block needs to be further analyzed. After obtaining the arrangement direction of the characters, establish the connection path of the characters. The path is connected in the order of the characters. For example, in a text block arranged in three lines, the characters in the first line are connected in order, and after a line break, they are connected to the first character in the second line, and so on. After the connection path is formed, call the offset data to sort the path. The sorting method is determined according to the arrangement order of the characters. If the text is left-aligned, it is sorted in ascending order of the horizontal coordinate values. If the text is in a vertical arrangement format, it is sorted in ascending order of the vertical coordinate values. During the sorting process, if it is found that the offset of a character jumps significantly, for example, the horizontal offset value of a character is greater than 50 pixels or the vertical offset value is less than 5 pixels, it may mean that the text has a non-standard arrangement or a paragraph break, and the sorting rule needs to be further adjusted to ensure the rationality of the character offset path sequence.
[0067] S103: Call the character offset path sequence, match the graphic-text combination label and the editing attribute status identifier, analyze the path change trend, filter the combination methods that conform to the structure, adjust the graphic-text association relationship, and obtain the initial graphic-text structure diagram;
[0068] First, analyze the path change trend and filter the combination methods that conform to the structure. The specific methods include counting the offset relationships of adjacent characters in the path. If the character offset values are within a stable range, for example, the fluctuation in the horizontal direction does not exceed 5 pixels, and the fluctuation in the vertical direction does not exceed 10 pixels, it can be determined that the text block is regularly arranged. If the offset values change greatly, the path matching method needs to be adjusted. In addition, for the relative positions of the characters and the images, adjust the graphic-text association relationship. For example, in a document with a picture caption, if the text is below the picture and the horizontal offset is stable within 40 pixels, it is considered that the text is associated with the picture. If the offset value of the text is large, such as exceeding 100 pixels, it may belong to an independent text block and the association relationship needs to be readjusted. After the association adjustment is completed, correct the final text path to ensure that the relative positions of the characters and the images are consistent. Finally, output the initial graphic-text structure diagram, in which the relative position relationships of all text blocks and images are matched and conform to the set alignment rules.
[0069] The specific steps of S2 are as follows:
[0070] S201: Call the graphic combination path in the initial graphic structure diagram, obtain the coordinate positions of the graphic elements in the current view, calculate the difference between the current view coordinates and the original path vector, establish the corresponding relationship between the two, and obtain the graphic coordinate comparison data;
[0071] First, parse the structure diagram data, extract the coordinate information of the graphic elements it contains, traverse the coordinate points of all graphic elements, record their positions in the original view, store these coordinate points as an ordered data set. Subsequently, obtain the coordinate positions of the graphic elements in the current view, that is, the element coordinates in the graphic combination presented on the current interface. Compare the current view coordinates with the original path coordinates, calculate the horizontal and vertical coordinate differences between the two, which are defined as the horizontal offset and the vertical offset. If the horizontal offset between the current view coordinates and the original path coordinates of an element is greater than 50 pixels and the vertical offset is less than 5 pixels, it is determined that the element has undergone a horizontal displacement. If the vertical offset is greater than 30 pixels and the horizontal offset is less than 5 pixels, it is considered that the element has undergone a vertical offset. After calculating the offset data of all elements, establish the corresponding relationship between the original path coordinates and the current view coordinates, and store all the calculated coordinate difference data in the coordinate comparison data set. If the offset values of some coordinates exceed a specific threshold, such as more than 100 pixels in the horizontal direction or more than 50 pixels in the vertical direction, further determine whether the data is abnormal and adjust it according to the data distribution. Finally, obtain the graphic coordinate comparison data.
[0072] S202: Based on the graphic coordinate comparison data, calculate the overall scaling and scrolling offset index, extract the offset angle and path distance difference between adjacent data points during the coordinate transformation process, analyze the path change trend, filter the change data of the path offset direction, and obtain the path offset calculation result;
[0073] The specific formula for the overall scaling and scrolling offset index is:
[0074] ;
[0075] Where, represents the overall scaling and scrolling offset index, represents the horizontal coordinate value of the th coordinate point in the current view, represents the horizontal coordinate value of the th coordinate point in the original view, represents the vertical coordinate value of the th coordinate point in the current view, represents the vertical coordinate value of the th coordinate point in the original view, represents the path distance of the th coordinate point in the original view, A small positive number to prevent the denominator from being zero, represents the total number of coordinate data points, represents the th scrolling displacement value in the current view, represents the th scrolling displacement value in the original view, represents the total number of scrolling displacement data points;
[0076] The formula is used to calculate the overall zoom magnification and scrolling displacement index , which involves two parts: graphic zooming / displacement and scrolling displacement. This formula combines actual monitoring data to quantitatively analyze the changes in graphic coordinates.
[0077] The first part calculates the zooming and displacement indices. Given the example data:
[0078] Assume there are data points, and their original and current horizontal and vertical coordinates are as follows: , , ;
[0079] The original distance can be obtained by calculating the Euclidean distance from each point to the origin. Here is just an example, and the calculation is simplified: (calculated by )
[0080] A small positive number to avoid division-by-zero errors.
[0081] Calculation process:
[0082] ;
[0083] ;
[0084] ;
[0085] ;
[0086] The second part calculates the scrolling displacement. Given the example data:
[0087] Assume scrolling data points, the original and current scrolling values: ,
[0088] Calculation process:
[0089] ;
[0090] Add the results of the two parts to obtain the total scaled scrolling offset metric :
[0091] ;
[0092] This result indicates that the overall graphic - text combination has experienced an average position offset of approximately 8.83% during the view change. This metric quantifies the magnitude of the element position change and provides a numerical basis for analyzing the stability of elements during view changes.
[0093] S203: Invoke the calculation result of path offset, analyze the structural changes of the combined path, identify the adjustment status of the graphic - text combination relationship, screen the feature points of the changed path, extract the corresponding offset patterns, and obtain the path offset identification set;
[0094] First, count the offset patterns of all elements in the path to determine whether there is an overall offset. If the horizontal or vertical offset values of all graphic - text elements tend to a certain fixed direction, such as all shifting 30 pixels to the right, it is considered that the overall path has shifted. If the offset values show a large uneven distribution, for example, some elements shift to the right while some elements shift to the left, there may be local adjustments. Subsequently, screen the feature points of the changed path, extract the nodes that have changed significantly compared to the original path, and record the offset patterns of these nodes. For example, if the offset direction of an element is left - aligned in the original view and becomes centered - aligned in the current view, the offset pattern of this node can be classified as an alignment - method adjustment. After all feature points are extracted, summarize the offset patterns to finally obtain the path offset identification set, which contains all the changed path nodes and their offset - pattern classification data.
[0095] The specific steps of S3 are as follows:
[0096] S301: Invoke the path offset identification set, extract the graphic - text combination indexes with path offsets, screen the corresponding graphic - text combination data, establish an index list, and obtain the graphic - text offset index set;
[0097] First, traverse the offset data stored in the path offset identifier set, obtain all the graphic-text combinations with position changes, record their index numbers in the initial structure diagram, extract the graphic-text combination data, that is, obtain the initial positions, current offset positions, and corresponding offsets of these graphic-text combinations. Subsequently, screen the graphic-text combinations that meet the offset characteristics, that is, analyze the offset amplitude. If the offset value of a certain graphic-text combination is within the range of 10 to 30 pixels, it is judged as a slight offset. If it exceeds 50 pixels, it is judged as a large offset. At the same time, check the offset direction. If the offset direction remains consistent in the horizontal or vertical direction, it is judged that the offset belongs to the overall displacement. If there are multiple direction changes in the offset direction, it is further subdivided into rotation, dislocation, or deformation offset. Then, establish an index list, that is, sort the screened graphic-text combinations according to the index numbers and store all the graphic indexes with offsets. After the index is established, record the graphic offset index set, which contains the offset status and index information of all graphic-text combinations.
[0098] S302: Based on the graphic offset index set, collect the timestamps, layer levels, and action tags in the user's editing operations, calculate the average trigger time interval, screen the action sequences that conform to the path offset law, analyze the fitness between the tags, identify the correlation relationships between actions, and obtain the action trigger matching data;
[0099] The specific formula for calculating the average trigger time interval is:
[0100] ;
[0101] where, represents the average trigger time interval of the th operation relative to the previous operation, represents the timestamp of the th operation in the current state, represents the th operation, represents the adjustment coefficient for the change in layer level, represents the layer level of the th layer in the current state, represents the th layer, represents the total number of operation records in the interface, represents the trigger path offset of the th operation in the current state, represents the trigger path offset of the th operation in the original state, represents the total number of path offset data points involved in the interface;
[0102] This formula is used to calculate the trigger spacing between adjacent operations and consists of two main parts. The first part calculates the mean time interval of all operations and introduces an influence factor for layer-level changes; the second part calculates the root mean square of the path offset changes to quantify the amplitude of the operation path changes.
[0103] Data monitoring and parameter acquisition;
[0104] Use an event listening tool to record the timestamps of the user's editing operations to obtain the operation time series .
[0105] Record all layer numbers in the editing interface And compare with the layer numbers of the previous interface .
[0106] Calculate the path offset through cursor trajectory data or interaction path , and compare with the previous state .
[0107] Weight parameter and adjustment coefficient setting;
[0108] Layer-level adjustment coefficient Value;
[0109] The setting range is , which is used to measure the impact of layer changes on the trigger spacing. The specific value is determined according to the complexity of layer adjustment. If the interface layer adjustment is less, the value is lower; if the layer adjustment is frequent, the value increases.
[0110] Sample quantity Value;
[0111] is the number of valid operations, is the number of valid path offset data points.
[0112] Specific calculation process (example);
[0113] Set the timestamp data of three operations (unit: millisecond):
[0114] ;
[0115] Calculate the adjacent time intervals:
[0116] ;
[0117] Set the layer numbers:
[0118] ;
[0119] Layer changes:
[0120] ;
[0121] Set , calculate the hierarchical influence factor:
[0122] ;
[0123] Calculate the average trigger time:
[0124] ;
[0125] Set the path offset data:
[0126] ;
[0127] ;
[0128] Calculate the sum of squares of path offsets:
[0129] ;
[0130] Calculate the root mean square:
[0131] ;
[0132] Finally, calculate the trigger spacing:
[0133] ;
[0134] This result shows that the average trigger spacing between adjacent operations is 401.67 milliseconds, reflecting the rhythm of the current user operations and the influence of interface hierarchical adjustment, providing a quantitative basis for interface response speed and path matching analysis.
[0135] S303: Invoke the action trigger matching data, construct a continuous operation node link, extract the trigger order within the graphic and text combination, filter the key nodes of path offset, and obtain the path offset trigger linked list;
[0136] Arrange all matching operations in chronological order, extract the trigger order within each graphic-text combination, count all the graphic-text combinations with offsets, and analyze the operation order before and after the offset. If a graphic-text combination first performs a "drag" operation and then a "resize" operation, then determine its operation order as scale after move. If the operation order of a graphic-text combination is "rotate" followed by "adjust angle", then record its operation order as the angle adjustment link. Then, screen the key nodes of the path offset, that is, find the operations that cause large offsets in the graphic-text combination. If the offset value of an operation is greater than 30 pixels, then determine it as a key offset node. If an operation triggers offsets in multiple graphic-text combinations simultaneously, then record this operation as a global impact node. After screening out all key operations, establish a path offset trigger linked list to store the key nodes and corresponding operation orders in all offset paths, and form a complete offset link data structure.
[0137] The specific steps of S4 are as follows:
[0138] S401: Invoke the path offset trigger linked list, extract consecutive operation nodes, obtain the corresponding audio-visual content types, screen the content data with synchronizable characteristics, and obtain the audio-visual type matching set;
[0139] First, traverse the operation sequence stored in the trigger linked list, identify the media content involved in each operation, determine whether it belongs to an audio or video element, and record the file format, duration, and playback status of these elements. For the content of audio-visual type, further screen whether it has synchronizable characteristics, that is, determine whether the audio and video can be aligned on the time axis. If the start time of an audio file has a fixed offset from the corresponding video file, such as within 0.5 seconds, then it is considered to have synchronizable characteristics. If the offset time exceeds 2 seconds, then its time relationship needs to be further analyzed. Subsequently, screen the content data that meets the synchronizable characteristics, that is, extract all audio-visual content that can be played synchronously, and establish an audio-visual matching index. Finally, obtain the audio-visual type matching set, which stores all audio-visual elements that meet the synchronizable characteristics and their associated operation nodes.
[0140] S402: Based on the audio-visual type matching set, calculate the rendering delay and the operation response time difference, extract the task loading timeliness data, calculate the synchronization difference between task loading and response, screen the task nodes that meet the synchronizable characteristics, and obtain the calculation result of the task synchronization difference;
[0141] First, record the actual loading time of the audio - video file during playback, and measure the time interval from when the user triggers an operation to when the audio - video content rendering is completed. If the loading time of the audio - video is less than 500 milliseconds, and the time interval from user operation to rendering is between 300 milliseconds and 1 second, then it is determined that the audio - video responds quickly. If the loading time exceeds 1 second, and the operation response time interval exceeds 2 seconds, then it is judged as a slow response. Then, extract the timeliness data of task loading, that is, traverse all audio - video playback tasks, count their loading durations, and calculate the synchronization difference between loading and response of the tasks, that is, the interval between when the user operation triggers playback and the actual playback time. If this interval is less than 1 second, then the task synchronization is considered good. If the interval exceeds 2 seconds, then timing compensation is required. After calculating the synchronization differences of all tasks, filter the task nodes that meet the synchronization characteristics, that is, find the tasks with synchronization errors within an acceptable range. Finally, obtain the calculation result of the task synchronization difference, where the synchronization error data of all audio - video tasks and the list of synchronizable tasks are stored.
[0142] S403: Call the calculation result of the task synchronization difference, judge the compliance of the channel multiplexing tolerance interval, extract the task indexes that meet the multiplexing conditions, filter the reusable task data, and obtain the set of reusable task indexes;
[0143] First, determine the time tolerance standard for channel multiplexing, that is, set the maximum tolerance value within the allowable range of task synchronization error. If the synchronization error of a certain task is between 0 and 500 milliseconds, then it is determined that it can perform channel multiplexing. If it exceeds 1 second, then it needs to be processed separately. Subsequently, extract the task indexes that meet the multiplexing conditions, that is, filter out all tasks with synchronization errors within the set range, and record their task numbers. Then, filter the reusable task data, that is, obtain all audio - video tasks that meet the tolerance standard, and analyze their multiplexing priorities. If the synchronization error of a certain task is less than 200 milliseconds and belongs to the same playback group, then it is preferentially multiplexed. If the synchronization error of a certain task is around 500 milliseconds but has the same start time as other tasks, then it can still be multiplexed. Finally, obtain the set of reusable task indexes, where the task numbers of all tasks that can perform channel multiplexing and their corresponding multiplexing parameters are stored.
[0144] The specific steps of S5 are as follows:
[0145] S501: Call the set of reusable task indexes, extract the graphic - text combination information, identify the positions of the corresponding elements in the editing interface, and filter all the matched combined elements to obtain the graphic - text combination positioning data;
[0146] First, traverse all task data stored in the task index set, filter out tasks involving a combination of graphics and text, extract their associated image and text elements, obtain the coordinate information of each graphics and text element, and store the position information of each element in the current editing interface. Subsequently, compare with the original graphics and text structure diagram to match the relative positions of the graphics and text elements. If the distance between a certain text block and the corresponding image remains within the range of 20 to 50 pixels, it is judged as a valid match. If it exceeds 100 pixels, it is marked as a potentially misaligned element. Then, filter all the matched combined elements, that is, eliminate the graphics and text combinations with no association or large matching deviations. Finally, establish the graphics and text combination positioning data, which stores all the matched graphics and text combinations and their coordinate information in the current editing interface.
[0147] S502: Based on the graphics and text combination positioning data, calculate the combined distance between the current element position and the original order, filter out text blocks that meet the logical distance requirements, extract the recombined sequence of the text blocks, call the element order matched by the combined distance data, and obtain the recombined sequence of the text combination;
[0148] First, obtain the current coordinates of all text blocks and calculate the displacement between them and the expected coordinates in the original order. If the displacement of a certain text block is between 10 and 30 pixels, it is judged that it basically maintains the original order. If it exceeds 50 pixels, further analyze its offset trend. Subsequently, filter out text blocks that meet the logical distance requirements, that is, eliminate text blocks with too large displacements, and only retain text blocks with small offsets and maintaining the relative order. Then, extract the recombined sequence of the text blocks, rearrange them according to the actual position order of the text blocks, and record the adjusted text order. If the order of a certain text block changes but the offset is within a reasonable range, add it to the recombined sequence. Finally, call the combined distance data, match the adjusted text element order, and store the recombined sequence of the text combination, which contains all text blocks that meet the logical order and their arrangement results.
[0149] S503: Call the recombined sequence of the text combination, re-establish the associated logical path, filter out the adjusted graphics and text combination structure, extract the sorting of the current document elements, and obtain the rearrangement result of the multimedia document structure;
[0150] First, obtain the association information between all text blocks and images, and adjust the graphics and text combination structure according to the recombined sequence. If the relative position offset between a certain text block and the original image is within 30 pixels, retain the original association. If it exceeds 50 pixels, reassign its corresponding image block. Subsequently, filter out the adjusted graphics and text combination structure, that is, eliminate the text-image combinations with unreasonable associations due to position changes, and rearrange all valid combinations. Then, extract the sorting information of the current document elements, that is, obtain the final arrangement order of all texts and images, and store it in the multimedia document structure data. Finally, obtain the rearrangement result of the multimedia document structure, which stores all the adjusted graphics and text combination orders and their final positions in the document.
[0151] Please refer to Figure 2 , a multimedia document online editing system, comprising:
[0152] The graphic and text structure analysis module obtains the pixel coordinates of the image boundary and the starting position of the corresponding text characters, calculates the coordinate difference in the alignment direction of the graphic and text blocks, generates a path sequence based on the offset direction of adjacent elements, matches the graphic and text combination tags and the editing attribute status identifiers, identifies the graphic and text combination relationship, and obtains the initial graphic and text structure diagram;
[0153] The view offset detection module calls the initial graphic and text structure diagram, detects the coordinate positions of the graphic and text elements in the view, compares them with the original path vector, calculates the zoom ratio, scroll displacement and offset angle, judges the change status of the combination relationship, and obtains the path offset identifier set;
[0154] The editing behavior analysis module calls the path offset identifier set, extracts the graphic and text combination index, collects the timestamps, layer levels and action tags of the user's editing operations, calculates the trigger interval and tag fitness of adjacent actions, constructs a continuous operation node link, and generates a path offset trigger linked list;
[0155] The task synchronization calculation module calls the path offset trigger linked list, analyzes the operations of the audio and video content, calculates the rendering delay and the operation response time difference, compares the synchronization difference between the loading timeliness and the response rhythm, judges the channel multiplexing tolerance interval, and generates a reusable task index set;
[0156] The document structure rearrangement module calls the reusable task index set, extracts the positions of the graphic and text elements in the editing interface, calculates the combination distance index from the original order, filters the text blocks with matching logical distances, reconstructs the graphic and text combination order and the associated logical path, and generates the multimedia document structure rearrangement result.
[0157] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An online editing method for multimedia documents, characterized in that, It includes the following steps: S1: Obtain the image boundary coordinates and the starting position of the text, extract the difference in combination with the graphic-text alignment direction, generate a path sequence based on the offset of adjacent elements, match the combined tags and the editing attribute status, and construct an initial graphic-text structure diagram; S2: Call the path of the initial structure diagram, compare the current graphic-text coordinates with the original path vector, calculate the offset angle and the path difference through the scaling ratio and the scrolling displacement, and generate a set of path offset identifiers; S3: Call the graphic-text combination index in the set of offset identifiers, collect the timestamp, layer level, and action tags during user editing, calculate the trigger interval and the tag fitting degree between actions, filter out continuous operation nodes, extract the trigger order within the graphic-text combination, and generate a path offset trigger linked list; S4: Call the audio-visual content types associated with the continuous operation nodes in the path offset trigger linked list, obtain the rendering delay and the response time difference, calculate the task synchronization deviation, determine whether it meets the channel multiplexing tolerance interval, and generate a set of reusable task indexes; S5: According to the set of reusable task indexes, locate the corresponding elements in the editing interface, compare their combined distance with the original order, reconstruct the logical order of the text blocks, and generate the rearrangement result of the multimedia document structure; The specific steps of S3 are as follows: S301: Call the set of path offset identifiers, extract the graphic-text combination indexes with path offsets, filter the corresponding graphic-text combination data, establish an index list, and obtain a set of graphic-text offset indexes; S302: Based on the set of graphic-text offset indexes, collect the timestamp, layer level, and action tags in the user editing operation, calculate the average trigger time interval, filter out the action sequences that conform to the path offset rule, analyze the fitting degree between tags, identify the association relationship between actions, and obtain action trigger matching data; S303: Call the action trigger matching data, construct a continuous operation node link, extract the trigger order within the graphic-text combination, filter out the key nodes of the path offset, and obtain a path offset trigger linked list; The specific steps of S4 are as follows: S401: Call the path offset trigger linked list, extract continuous operation nodes, obtain the corresponding audio-visual content types, filter out the content data with synchronizable characteristics, and obtain a set of audio-visual type matches; S402: Based on the set of audio-visual type matches, calculate the rendering delay and the operation response time difference, extract the task loading time efficiency data, calculate the synchronization difference between task loading and response, filter out the task nodes that meet the synchronization characteristics, and obtain the calculation result of the task synchronization difference; S403: Call the calculation result of the task synchronization difference, judge the compliance with the channel multiplexing tolerance interval, extract the task indexes that meet the multiplexing conditions, filter out the reusable task data, and obtain a set of reusable task indexes; The specific steps of S5 are as follows: S501: Call the set of reusable task indexes, extract the graphic-text combination information, identify the positions of the corresponding elements in the editing interface, filter out all the matched combined elements, and obtain the graphic-text combination positioning data; S502: Based on the combined graphic and text positioning data, calculate the combined distance between the current element position and the original order, filter out text blocks that meet the logical distance requirements, extract the recombined sequence of text blocks, call the element order matching the combined distance data, and obtain the text combination reconstruction sequence; S503: Call the text combination reconstruction sequence, re - establish the associated logical path, filter out the adjusted graphic and text combination structure, extract the current document element sorting, and obtain the multimedia document structure rearrangement result.
2. The online editing method of a multimedia document according to claim 1, wherein, The initial graphic and text structure diagram includes image boundary coordinates, text character positions, path direction information, combined label types, and editing attribute statuses. The path offset identification set includes offset angle values, path distance differences, scaling ratio parameters, and scrolling displacement amounts. The path offset trigger linked list includes operation time nodes, layer - level identifiers, action type tags, and trigger order information. The reusable task index set includes audio - video content types, rendering delay parameters, response time difference indicators, and channel multiplexing tolerance ranges. The multimedia document structure rearrangement result includes element combined distances, text block order relationships, associated logical paths, and interface positioning positions.
3. The online editing method of a multimedia document according to claim 1, wherein The specific steps of S1 are as follows: S101: Obtain the pixel coordinates of the image boundary, match the starting position of the text character, calculate the offset between the two, extract the coordinate differences in the horizontal and vertical directions, establish a coordinate matching relationship, and obtain the pixel - text matching offset. S102: Based on the pixel - text matching offset, extract the offset directions of adjacent characters, calculate the displacement difference, filter out data that meet the text block alignment direction, generate a character connection path, and call the offset data for path sorting to obtain the character offset path sequence. S103: Call the character offset path sequence, match the graphic and text combination label with the editing attribute status identifier, analyze the path change trend, filter out the combination methods that meet the structure, adjust the graphic and text association relationship, and obtain the initial graphic and text structure diagram.
4. The online editing method of a multimedia document according to claim 1, wherein The specific steps of S2 are as follows: S201: Call the graphic and text combination path in the initial graphic and text structure diagram, obtain the coordinate positions of the graphic and text elements in the current view, calculate the difference between the current view coordinates and the original path vector, establish the corresponding relationship between the two, and obtain the graphic and text coordinate comparison data. S202: Based on the graphic and text coordinate comparison data, calculate the overall scaling and scrolling offset index, extract the offset angle and path distance difference between adjacent data points during the coordinate transformation process, analyze the path change trend, filter out the changing data in the path offset direction, and obtain the path offset calculation result. S203: Call the path offset calculation result, analyze the structural changes of the combined path, identify the adjustment status of the graphic and text combination relationship, filter out the feature points of the changing path, and extract the corresponding offset mode to obtain the path offset identification set.
5. The online editing method of a multimedia document according to claim 4, wherein The specific formula for the overall scaling and scrolling offset index is as follows: ; Among them, represents the overall scaling scroll offset index, represents the horizontal coordinate value of the th coordinate point in the current view, represents the horizontal coordinate value of the th coordinate point in the original view, represents the vertical coordinate value of the th coordinate point in the current view, represents the vertical coordinate value of the th coordinate point in the original view, represents the path distance of the th coordinate point in the original view, is a tiny positive number to prevent the denominator from being zero, represents the total number of coordinate data points, represents the th scroll displacement value in the current view, represents the th scroll displacement value in the original view, represents the total number of scroll displacement data points.
6. The online editing method of a multimedia document according to claim 1, characterized in that, The specific formula for the average trigger time interval is as follows: ; Among them, represents the average trigger time interval of the th operation relative to the previous operation, represents the timestamp of the th operation in the current state, represents the th operation timestamp, represents the adjustment coefficient for layer level changes, represents the layer level of the th layer in the current state, represents the th layer level, represents the number of operations recorded in the interface, represents the trigger path offset of the th operation in the current state, represents the trigger path offset of the th operation in the original state, represents the total number of path offset data points involved in the interface.
7. An online editing system for multimedia documents, characterized in that According to any one of claims 1 - 6, a method for online editing of multimedia documents, the system includes: The graphic-text structure analysis module obtains the pixel coordinates of the image boundary and the starting positions of the corresponding text characters, calculates the coordinate differences in the alignment direction of the graphic-text blocks, generates a path sequence based on the offset directions of adjacent elements, matches the graphic-text combination tags and the editing attribute status identifiers, identifies the graphic-text combination relationship, and obtains the initial graphic-text structure diagram; The view offset detection module calls the initial graphic-text structure diagram, detects the coordinate positions of the graphic-text elements in the view, compares them with the original path vectors, calculates the zoom ratio, scrolling displacement, and offset angle, judges the change status of the combination relationship, and obtains the path offset identifier set; The editing behavior analysis module calls the path offset identifier set, extracts the graphic-text combination index, collects the timestamps, layer levels, and action tags of the user's editing operations, calculates the trigger intervals and tag fitting degrees of adjacent actions, constructs a continuous operation node link, and generates a path offset trigger linked list; The task synchronization calculation module calls the path offset trigger linked list, analyzes the operations on the audio-visual content, calculates the rendering delay and the operation response time difference, compares the synchronization difference between the loading efficiency and the response rhythm, judges the channel multiplexing tolerance interval, and generates a reusable task index set; The document structure rearrangement module calls the reusable task index set, extracts the positions of the graphic-text elements in the editing interface, calculates the combination distance index from the original order, filters the text blocks with matching logical distances, reconstructs the graphic-text combination order and the associated logical path, and generates the multimedia document structure rearrangement result.
Citation Information
Patent Citations
Multimedia data processing method, device and equipment and readable storage medium
CN114257843A
Document editing method and device, medium and electronic equipment
CN117010336A