Multimedia document online editing method and system
By constructing the initial structure diagram of the graphics and text and analyzing the editing behavior, the problem of insufficient dynamic analysis ability of the graphics and text combination relationship in the existing technology is solved, and the structural stability and interactive consistency of multimedia documents in complex editing scenarios is achieved.
Patent Information
- Application Number
- CN202510526981.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-25
AI Technical Summary
When dealing with the combination relationship of graphic and text, the prior art lacks the ability to dynamically analyze the relative coordinates and offset paths between graphic and text, which leads to difficulties in dislocation and reconstruction of images and texts in documents. In scenarios such as multi-user collaboration, cross-terminal editing, and complex graphic and text mixing, it is difficult to ensure the structural stability and interactive consistency of the content.
By obtaining the image boundary coordinates and text starting positions, extracting the difference value in combination with the alignment direction of the text, generating a path sequence, matching the combination labels and editing attribute status, and building an initial structure diagram of the text. Then, through the path offset identification set and trigger link list, the synchronization of user editing behavior and audio and video content is analyzed, the task synchronization deviation is calculated, the reusable task index set is generated, and the text block logical sequence is finally reconstructed, and the multimedia document structure rearrangement results are generated.
It realizes dynamic analysis of the graphic and text combination relationship and effective modeling of editing behavior, improves the timing consistency of content reorganization, ensures the consistency and precise positioning of structure reordering, and enhances the stability and response efficiency of multimedia documents in complex editing scenarios.
Smart Images

Figure CN120068817A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of document processing, and in particular, to a method and system for online editing of multimedia documents. Background Art
[0002] The technical field of document processing includes the entire process of document creation, editing, management, retrieval, and storage, covering comprehensive processing methods for multimedia content such as text, images, audio, and video. The core content of this technical field lies in realizing the unified access, structured processing, and visual operation of multi-source heterogeneous data, and supporting key links such as collaborative work, permission control, and version management. With the development of information technology, document processing technology has gradually shifted from local processing to online and intelligent processing, and is widely used in fields such as office automation, content management systems, and digital archive management. It has the ability to efficiently process large-scale document data and support remote real-time collaboration, reflecting a strong dependence on the integration of data organization and content services.
[0003] Among them, the method for online editing of multimedia documents refers to a technical method for complex documents containing elements such as text, pictures, audio, and video, which provides an editing interface through a network platform to support remote real-time collaborative editing. This method usually covers multiple technical matters such as format parsing, unified rendering, logical structure recognition, cross-terminal compatibility processing, and cloud data synchronization of multimedia content. The specific methods include marking and parsing multiple types of content based on a predefined document structure model, using a web editing kernel to perform online loading and editing control of multimedia elements, and completing real-time synchronization of document status and transmission and storage of modified data through a communication protocol between the server and the client. Such methods generally rely on document content structure modeling, combined with a real-time network communication mechanism and a browser visualization interface to achieve concurrent access and editing by multiple users.
[0004] In the prior art, when dealing with the combination relationship between graphics and text, it relies on a static structure model for position determination, lacking the ability to dynamically analyze the relative coordinates and offset paths between graphics and text, resulting in misalignment and difficulty in reconstruction of images and text in the document. In the face of high-frequency and multi-dimensional operation inputs from users, an effective modeling of operation timing and trigger chains has not been formed, making it difficult to restore the front and back causality in the operation chain, resulting in problems such as chaotic structure update logic and disordered editing content sequence. In the scenario of task concurrency and audio-video mixed content, the current mechanism lacks quantitative control over the synchronization degree of content loading and operation response, easily causing editing stagnation or task blocking. In the structure adjustment stage, existing methods mostly perform simple rearrangement based on the content presentation order, lacking fine calculation of logical paths and structure spacing, resulting in problems such as logical breakage and disordered combination relationships after document structure rearrangement. These deficiencies are particularly prominent in application environments such as multi-user collaboration, cross-terminal editing, and complex graphic-text mixed layout, directly affecting the structural stability of document content and the interaction coherence during use. Summary of the Invention
[0005] The object of the present invention is to solve the deficiencies existing in the prior art, and a method and system for online editing of multimedia documents are proposed.
[0006] To achieve the above object, the present invention adopts the following technical solutions: A method for online editing of multimedia documents, comprising the following steps: S1: Obtain the image boundary coordinates and the starting position of the text, extract the difference in combination with the graphic-text alignment direction, generate a path sequence based on the offset of adjacent elements, match the combined tags and the editing attribute status, and construct an initial graphic-text structure diagram; S2: Call the path of the initial structure diagram, compare the current graphic-text coordinates with the original path vector, calculate the offset angle and the path difference through the scaling ratio and the scrolling displacement, and generate a path offset identification set; S3: Call the graphic-text combination index in the offset identification set, collect the timestamp, layer level, and action tags during user editing, calculate the trigger interval between actions and the tag fitting degree, screen continuous operation nodes, extract the trigger order within the graphic-text combination, and generate a path offset trigger linked list; S4: Call the audio-visual content types associated with the continuous operation nodes in the path offset trigger linked list, obtain the rendering delay and the response time difference, calculate the task synchronization deviation, determine whether it meets the channel multiplexing tolerance interval, and generate a reusable task index set; S5: According to the reusable task index set, locate the corresponding elements on the editing interface, compare their combined distance with the original order, reconstruct the logical order of the text blocks, and generate a multimedia document structure rearrangement result.
[0007] As a further solution of the present invention, the initial graphic-text structure diagram includes image boundary coordinates, text character positions, path direction information, combined tag types, and editing attribute statuses. The path offset identification set includes offset angle values, path distance differences, scaling ratio parameters, and scrolling displacement amounts. The path offset trigger linked list includes operation time nodes, layer level identifiers, action type tags, and trigger order information. The reusable task index set includes audio-visual content types, rendering delay parameters, response time difference indicators, and channel multiplexing tolerance ranges. The multimedia document structure rearrangement result includes element combination distances, text block order relationships, associated logical paths, and interface positioning positions.
[0008] As a further solution of the present invention, the specific steps of S1 are: S101: Obtain the image boundary pixel coordinates, match the starting position of the text characters, calculate the offset between the two, extract the coordinate differences in the horizontal and vertical directions, establish a coordinate matching relationship, and obtain the pixel-text matching offset; S102: Based on the pixel text matching offset, extract the offset directions of adjacent characters, calculate the displacement difference, filter the data that conforms to the text block alignment direction, generate a character connection path, and call the offset data for path sorting to obtain a character offset path sequence; S103: Call the character offset path sequence, match the graphic-text combination tag and the editing attribute status identifier, analyze the path change trend, filter the combination methods that conform to the structure, adjust the graphic-text association relationship, and obtain the initial graphic-text structure diagram.
[0009] As a further solution of the present invention, the specific steps of S2 are as follows: S201: Call the graphic-text combination path in the initial graphic-text structure diagram, obtain the coordinate positions of the graphic-text elements in the current view, calculate the difference between the current view coordinates and the original path vector, establish the corresponding relationship between the two, and obtain the graphic-text coordinate comparison data; S202: Based on the graphic-text coordinate comparison data, calculate the overall scaling and scrolling offset index, extract the offset angle and the path distance difference between adjacent data points during the coordinate transformation process, analyze the path change trend, filter the change data of the path offset direction, and obtain the path offset calculation result; S203: Call the path offset calculation result, analyze the structural changes of the combination path, identify the adjustment status of the graphic-text combination relationship, filter the feature points of the changed path, and extract the corresponding offset mode to obtain a path offset identifier set.
[0010] As a further solution of the present invention, the specific formula for the overall scaling and scrolling offset index is as follows: ; Wherein, represents the overall scaling and scrolling offset index, represents the horizontal coordinate value of the th coordinate point in the current view, represents the horizontal coordinate value of the th coordinate point in the original view, represents the vertical coordinate value of the th coordinate point in the current view, represents the vertical coordinate value of the th coordinate point in the original view, represents the path distance of the th coordinate point in the original view, is a tiny positive number to prevent the denominator from being zero, represents the total number of coordinate data points, represents the th scrolling displacement value in the current view, represents the th scrolling displacement value in the original view, Represents the total number of rolling displacement data points.
[0011] As a further solution of the present invention, the specific steps of S3 are as follows: S301: Invoke the path offset identification set, extract the graphic-text combination indexes with path offsets, screen the corresponding graphic-text combination data, establish an index list, and obtain a graphic-text offset index set; S302: Based on the graphic-text offset index set, collect the timestamps, layer levels, and action tags in the user editing operation, calculate the average trigger time interval, screen the action sequences that conform to the path offset rule, analyze the fitting degree between the tags, identify the correlation relationship between actions, and obtain action trigger matching data; S303: Invoke the action trigger matching data, construct a continuous operation node link, extract the trigger order within the graphic-text combination, screen the key nodes with path offsets, and obtain a path offset trigger linked list.
[0012] As a further solution of the present invention, the specific formula for calculating the average trigger time interval is as follows: ; Wherein, represents the average trigger time interval of the th operation relative to the previous operation, represents the timestamp of the th operation in the current state, represents the timestamp of the th operation, represents the adjustment coefficient of the layer level change, represents the layer level of the th layer in the current state, represents the layer level of the th layer, represents the number of operations recorded in the interface, represents the trigger path offset of the th operation in the current state, represents the trigger path offset of the th operation in the original state, represents the total number of path offset data points involved in the interface.
[0013] As a further solution of the present invention, the specific steps of S4 are as follows: S401: Invoke the path offset trigger linked list, extract continuous operation nodes, obtain the corresponding audio-visual content types, screen the content data with synchronizable characteristics, and obtain an audio-visual type matching set; S402: Based on the audio - video type matching set, calculate the rendering delay and the operation response time difference, extract the task loading timeliness data, calculate the synchronization difference between task loading and response, filter the task nodes that meet the synchronization characteristics, and obtain the task synchronization difference calculation result; S403: Invoke the task synchronization difference calculation result, judge the compliance with the channel multiplexing tolerance interval, extract the task indexes that meet the multiplexing conditions, filter the reusable task data, and obtain the reusable task index set.
[0014] As a further solution of the present invention, the specific steps of S5 are as follows: S501: Invoke the reusable task index set, extract the graphic - text combination information, identify the corresponding element positions in the editing interface, filter all the matched combination elements, and obtain the graphic - text combination positioning data; S502: Based on the graphic - text combination positioning data, calculate the combination distance between the current element position and the original order, filter the text blocks that meet the logical distance requirements, extract the recombination sequence of the text blocks, and invoke the element order matched with the combination distance data to obtain the text combination reconstruction sequence; S503: Invoke the text combination reconstruction sequence, re - establish the associated logical path, filter the adjusted graphic - text combination structure, extract the current document element sorting, and obtain the multimedia document structure rearrangement result.
[0015] A multimedia document online editing system includes: The graphic - text structure analysis module obtains the image boundary pixel coordinates and the corresponding text character start positions, calculates the coordinate difference in the alignment direction of the graphic - text blocks, generates a path sequence according to the offset direction of adjacent elements, matches the graphic - text combination tags and the editing attribute status identifiers, identifies the graphic - text combination relationship, and obtains the initial graphic - text structure diagram; The view offset detection module invokes the initial graphic - text structure diagram, detects the coordinate positions of the graphic - text elements in the view, compares with the original path vector, calculates the zoom ratio, scroll displacement and offset angle, and judges the change state of the combination relationship to obtain the path offset identifier set; The editing behavior analysis module invokes the path offset identifier set, extracts the graphic - text combination indexes, collects the timestamps, layer levels and action tags of the user's editing operations, calculates the trigger interval and tag fitting degree of adjacent actions, constructs a continuous operation node link, and generates a path offset trigger linked list; The task synchronization calculation module invokes the path offset trigger linked list, analyzes the operations of the audio - video content, calculates the rendering delay and the operation response time difference, compares the synchronization difference between the loading timeliness and the response rhythm, judges the channel multiplexing tolerance interval, and generates a reusable task index set; The document structure rearrangement module calls the reusable task index set, extracts the positions of graphic and text elements in the editing interface, calculates the combined distance index from the original order, filters text blocks with matching logical distances, reconstructs the graphic and text combination order and associated logical paths, and generates the result of multimedia document structure rearrangement.
[0016] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In the present invention, a logical structure is constructed through the difference between graphic and text coordinates and the offset path, the change of the combination state is accurately judged by combining the angle and distance changes, the time series and label features in the editing behavior are collected, an operation link is established to improve the timing consistency of content recombination, the rhythm control of content loading is realized by calculating the task synchronization rhythm through the delay and response difference, the combination order is reconstructed according to the logical distance, and the coherence and precise positioning after structure rearrangement are maintained, enhancing the stability and response efficiency of multimedia documents in complex editing scenarios. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a schematic diagram of the step flow of the present invention; Figure 2 It is a system module diagram of the present invention. Detailed Embodiments
[0019] The following will describe the technical solutions in the present invention with reference to the drawings.
[0020] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "example" in the present invention should not be construed as more preferred or more advantageous than other embodiments or design solutions. Exactly speaking, the use of the word "example" is intended to present concepts in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.
[0021] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, their intended meanings are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, their intended meanings are the same.
[0022] In the embodiments of the present invention, sometimes subscripts such as W 1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meanings they express are the same.
[0023] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0024] Please refer to Figure 1 , a method for online editing of multimedia documents, comprising the following steps: S1: Obtain the pixel coordinates of the image boundary and the starting position of the corresponding text characters, extract the coordinate difference in combination with the alignment direction of the graphic-text block, generate a path sequence according to the offset direction between adjacent elements, and call the graphic-text combination tag and the editing attribute status identifier for combined matching to obtain the initial graphic-text structure diagram; S2: Call the graphic-text combination path in the initial graphic-text structure diagram, compare the coordinate positions of the graphic-text elements in the current view with the original path vector, calculate the offset angle and the path distance difference through the zoom ratio and the scroll displacement, and identify the change state of the combination relationship to obtain the path offset identifier set; S3: Call the graphic-text combination index with path offset in the path offset identifier set, collect the timestamp, layer level, and action label in the user editing operation, calculate the trigger interval and label fitting degree between adjacent actions, form a continuous operation node link, extract the trigger order within the graphic-text combination, and generate a path offset trigger linked list; S4: Call the audio-visual content types associated with the continuous operation nodes in the path offset trigger linked list, obtain the rendering delay and the operation response time difference, calculate the synchronization difference between the loading timeliness and the response rhythm of the current task, and judge the compliance of the channel multiplexing tolerance interval to generate a reusable task index set; S5: According to the graphic-text combination information in the reusable task index set, identify the corresponding element positions in the editing interface, compare the combination distance between the positions and the original order, select the text block corresponding to the logical distance to reconstruct the combination order, re-establish the associated logical path, and generate the multimedia document structure rearrangement result.
[0025] The initial graphic-text structure diagram includes image boundary coordinates, text character positions, path direction information, combination label types, and editing attribute statuses. The path offset identifier set includes offset angle values, path distance differences, zoom ratio parameters, and scroll displacement amounts. The path offset trigger linked list includes operation time nodes, layer level identifiers, action type labels, and trigger order information. The reusable task index set includes audio-visual content types, rendering delay parameters, response time difference indicators, and channel multiplexing tolerance ranges. The multimedia document structure rearrangement result includes element combination distances, text block order relationships, associated logical paths, and interface positioning positions.
[0026] The specific steps of S1 are as follows: S101: Obtain the pixel coordinates of the image boundary, match the starting position of the text characters, calculate the offset between the two, extract the coordinate differences in the horizontal and vertical directions, establish a coordinate matching relationship, and obtain the pixel-text matching offset; It is necessary to preprocess the input image first, including grayscale conversion and binarization, in order to highlight the text and boundary information in the image. On this basis, by traversing the pixels of the image, the areas with significant brightness changes are screened as possible boundary areas, and further, the contour lines of the image are extracted through edge detection methods. After obtaining the boundary pixel coordinates, they are stored as an ordered coordinate set. At the same time, the starting position of the text characters is located. Using optical character recognition (OCR) technology, the characters in the image are scanned and the upper-left coordinate points of each character are output to extract the starting positions of all characters. Subsequently, the horizontal and vertical coordinate deviations between the image boundary pixel coordinates and the starting positions of the text characters are calculated. The horizontal deviation value represents the position difference of the characters in the horizontal direction, and the vertical deviation value represents the position difference of the characters in the vertical direction. By traversing the deviation values of all characters, the distribution range of the deviation values is statistically analyzed, and the values with a relatively small change range are screened to determine the matching offset of the pixel text. For example, if the horizontal offset value of the character starting coordinates is stable between 30 and 35 pixels, and the vertical offset value is between 5 and 10 pixels, it can be determined that the pixel-text offset of this image is mainly concentrated in these two ranges. If the offset values of some characters are far from this range, it is necessary to judge whether they are abnormal data and further adjust the matching rules. Finally, the matching relationship between the pixel coordinates and the text coordinates is established to determine the matching offset.
[0027] S102: Based on the pixel-text matching offset, extract the offset directions of adjacent characters, calculate the displacement difference, screen the data that conforms to the text block alignment direction, generate a character connection path, and call the offset data to perform path sorting to obtain a character offset path sequence; Extract the offset directions of adjacent characters. First, arrange the characters in the order from left to right and from top to bottom, calculate the coordinate changes between every two adjacent characters, count the horizontal and vertical offset values of the characters, and filter the data that conforms to the alignment direction of the text block. For example, if the text is arranged in standard left alignment, the horizontal offset values of all adjacent characters should be within a certain range, such as between 20 and 25 pixels, and the vertical offset value is close to zero, indicating that the text is regularly arranged horizontally. If the vertical offset value is large, it indicates that there may be line spacing and further analysis of the text block structure is required. After obtaining the arrangement direction of the characters, establish the connection path of the characters. The path is connected in the order of the characters. For example, in a text block arranged in three lines, the characters in the first line are connected in order, and after a line break, they are connected to the first character in the second line, and so on. After the connection path is formed, call the offset data to sort the path. The sorting method is determined according to the arrangement order of the characters. If the text is left-aligned, it is sorted in ascending order of the horizontal coordinate values. If the text is in a vertical arrangement format, it is sorted in ascending order of the vertical coordinate values. During the sorting process, if it is found that the offset of a character jumps significantly, for example, the horizontal offset value of a character is greater than 50 pixels or the vertical offset value is less than 5 pixels, it may mean that the text has a non-standard arrangement or a paragraph break, and the sorting rule needs to be further adjusted to ensure the rationality of the character offset path sequence.
[0028] S103: Call the character offset path sequence, match the graphic-text combination label and the editing attribute status identifier, analyze the path change trend, filter the combination methods that conform to the structure, adjust the graphic-text association relationship, and obtain the initial graphic-text structure diagram; First, analyze the path change trend and filter the combination methods that conform to the structure. The specific methods include counting the offset relationship between adjacent characters in the path. If the character offset value is within a stable range, for example, the fluctuation in the horizontal direction does not exceed 5 pixels, and the fluctuation in the vertical direction does not exceed 10 pixels, it can be determined that the text block is regularly arranged. If the offset value changes greatly, the path matching method needs to be adjusted. In addition, for the relative position of the character and the image, adjust the graphic-text association relationship. For example, in a document with a picture caption, if the text is below the picture and the horizontal offset is stable within 40 pixels, it is considered that the text is associated with the picture. If the offset value of the text is large, such as more than 100 pixels, it may belong to an independent text block and the association relationship needs to be readjusted. After the association adjustment is completed, correct the final text path to ensure that the relative position of the character and the image remains consistent, and finally output the initial graphic-text structure diagram, in which the relative position relationship between all text blocks and images has been matched and conforms to the set alignment rules.
[0029] The specific steps of S2 are as follows: S201: Call the graphic combination path in the initial graphic structure diagram, obtain the coordinate positions of the graphic elements in the current view, calculate the difference between the current view coordinates and the original path vector, establish the corresponding relationship between the two, and obtain the graphic coordinate comparison data; First, parse the structure diagram data, extract the coordinate information of the graphic elements contained therein, traverse the coordinate points of all graphic elements, record their positions in the original view, store these coordinate points as an ordered data set. Subsequently, obtain the coordinate positions of the graphic elements in the current view, that is, the element coordinates in the graphic combination presented on the current interface. Compare the current view coordinates with the original path coordinates, calculate the horizontal and vertical coordinate differences between the two, which are defined as the horizontal offset and the vertical offset. If the horizontal offset between the current view coordinates and the original path coordinates of an element is greater than 50 pixels and the vertical offset is less than 5 pixels, it is determined that the element has undergone a horizontal displacement. If the vertical offset is greater than 30 pixels and the horizontal offset is less than 5 pixels, it is considered that the element has undergone a vertical offset. After calculating the offset data of all elements, establish the corresponding relationship between the original path coordinates and the current view coordinates, and store all the calculated coordinate difference data into the coordinate comparison data set. If the offset values of some coordinates exceed a specific threshold, such as more than 100 pixels in the horizontal direction or more than 50 pixels in the vertical direction, further determine whether the data is abnormal and adjust it according to the data distribution. Finally, obtain the graphic coordinate comparison data.
[0030] S202: Based on the graphic coordinate comparison data, calculate the overall scaling and scrolling offset index, extract the offset angle and path distance difference between adjacent data points during the coordinate transformation process, analyze the path change trend, screen the change data of the path offset direction, and obtain the path offset calculation result; The specific formula for the overall scaling and scrolling offset index is: ; Among them, represents the overall scaling and scrolling offset index, represents the horizontal coordinate value of the th coordinate point in the current view, represents the horizontal coordinate value of the th coordinate point in the original view, represents the vertical coordinate value of the th coordinate point in the current view, represents the vertical coordinate value of the th coordinate point in the original view, represents the path distance of the th coordinate point in the original view, is a tiny positive number to prevent the denominator from being zero, represents the total number of coordinate data points, Represents the th scrolling displacement value in the current view, represents the th scrolling displacement value in the original view, represents the total number of scrolling displacement data points; The formula is used to calculate the overall zoom magnification and scrolling displacement index , which involves two parts: graphic scaling / displacement and scrolling displacement. This formula combines actual monitoring data to quantitatively analyze the changes in graphic coordinates.
[0031] The first part calculates the scaling and displacement index, given example data: Assume there are data points, and their original and current horizontal and vertical coordinates are as follows: , , ; The original distance can be obtained by calculating the Euclidean distance from each point to the origin. Here is just an example, and the calculation is simplified: (calculated by ); A small positive number to avoid division by zero error.
[0032] Calculation process: ; ; ; ; The second part calculates the scrolling displacement, given example data: Assume scrolling data points, original and current scrolling values: ,
[0033] Calculation process: ; Add the results of the two parts to get the total scaling and scrolling offset index : ; This result indicates that the overall graphic combination has experienced an average position offset of about 8.83% during the view change. This index quantifies the change range of the element position and provides a numerical basis for analyzing the stability of elements during view change.
[0034] S203: Call the path offset calculation result, analyze the structural changes of the combined path, identify the adjustment status of the graphic-text combination relationship, screen the feature points of the changed path, extract the corresponding offset patterns, and obtain the path offset identification set; First, count the offset patterns of all elements in the path to determine whether there is an overall offset. If the horizontal or vertical offset values of all graphic-text elements tend to a certain fixed direction, such as all shifting 30 pixels to the right, it is considered that the overall path has shifted. If the offset values show a large uneven distribution, for example, some elements shift to the right while some shift to the left, there may be local adjustments. Subsequently, screen the feature points of the changed path, extract the nodes that have changed significantly compared to the original path, and record the offset patterns of these nodes. For example, if the offset direction of an element is left-aligned in the original view and becomes centered-aligned in the current view, the offset pattern of this node can be classified as an alignment adjustment. After all feature points are extracted, summarize the offset patterns, and finally obtain the path offset identification set, which contains all the changed path nodes and their offset pattern classification data.
[0035] The specific steps of S3 are as follows: S301: Call the path offset identification set, extract the graphic-text combination indexes with path offsets, screen the corresponding graphic-text combination data, establish an index list, and obtain the graphic-text offset index set; First, traverse the offset data stored in the path offset identification set, obtain all the graphic-text combinations with position changes, and record their index numbers in the initial structure diagram. Extract the graphic-text combination data, that is, obtain the initial positions, current offset positions, and corresponding offset amounts of these graphic-text combinations. Subsequently, screen the graphic-text combinations that meet the offset characteristics, that is, analyze the offset amplitude. If the offset value of a graphic-text combination is in the range of 10 to 30 pixels, it is judged as a slight offset. If it exceeds 50 pixels, it is judged as a large offset. At the same time, check the offset direction. If the offset direction remains consistent in the horizontal or vertical direction, it is judged that the offset belongs to an overall displacement. If the offset direction shows multiple direction changes, it is further subdivided into rotation, dislocation, or deformation offsets. Then, establish an index list, that is, sort the screened graphic-text combinations according to the index numbers and store all the graphic indexes with offsets. After the index is established, record the graphic-text offset index set, which contains the offset status and index information of all graphic-text combinations.
[0036] S302: Based on the graphic-text offset index set, collect the timestamps, layer levels, and action tags in the user's editing operations, calculate the average trigger time interval, screen the action sequences that conform to the path offset rules, analyze the fitness between the tags, identify the correlation relationships between the actions, and obtain the action trigger matching data; The specific formula for the average trigger time interval is: ; Among them, represents the average trigger time interval of the th operation relative to the previous operation, represents the timestamp of the th operation in the current state, represents the th operation timestamp, represents the adjustment coefficient for layer level changes, represents the layer level of the th layer in the current state, represents the th layer level, represents the number of operations recorded in the interface, represents the trigger path offset of the th operation in the current state, represents the trigger path offset of the th operation in the original state, represents the total number of path offset data points involved in the interface; This formula is used to calculate the trigger spacing between adjacent operations and consists of two main parts. The first part calculates the mean of the time intervals for all operations and introduces the influence factor of layer level changes; the second part calculates the root mean square of the path offset changes to quantify the amplitude of the operation path changes.
[0037] Data monitoring and parameter acquisition; Use an event listening tool to record the timestamps of the user's editing operations to obtain the operation time series .
[0038] Record all level numbers in the editing interface and compare with the level numbers of the previous interface .
[0039] Calculate the path offset through cursor trajectory data or interaction path , and compare with the previous state .
[0040] Weight parameter and adjustment coefficient setting; Layer level adjustment coefficient value; The setting range is , which is used to measure the impact of layer level changes on the trigger spacing. The specific value is determined according to the complexity of the layer level adjustment. If the interface layer level adjustment is less, the value is lower; if the layer level adjustment is frequent, the value increases.
[0041] Sample quantity value; is the number of valid operations, It is the number of valid path offset data points.
[0042] Specific calculation process (numerical example); Set the timestamp data for three operations (unit: millisecond): ; Calculate the adjacent time intervals: ; Set the hierarchy number: ; Hierarchy change: ; Set , and calculate the hierarchy influence factor: ; Calculate the average trigger time: ; Set the path offset data: ; ; Calculate the sum of squares of path offsets: ; Calculate the root mean square: ; Finally, calculate the trigger spacing: ; This result indicates that the average trigger spacing between adjacent operations is 401.67 milliseconds, reflecting the rhythm of the current user operations and the influence of interface hierarchy adjustment, providing a quantitative basis for interface response speed and path matching analysis.
[0043] S303: Invoke the action trigger matching data, construct a continuous operation node link, extract the trigger order within the graphic and text combination, filter the key nodes of path offset, and obtain the path offset trigger linked list; Arrange all matching operations in chronological order, extract the trigger order within each graphic-text combination, count all graphic-text combinations with offsets, and analyze the operation order before and after the offset. If a graphic-text combination first performs a "drag" operation and then a "resize" operation, then determine its operation order as scale after move. If the operation order of a graphic-text combination is "rotate" followed by "adjust angle", then record its operation order as the angle adjustment link. Then, screen the key nodes of the path offset, that is, find the operations that cause significant offsets in the graphic-text combinations. If the offset value of an operation is greater than 30 pixels, then determine it as a key offset node. If an operation triggers offsets in multiple graphic-text combinations simultaneously, then record this operation as a global impact node. After screening out all key operations, establish a path offset trigger linked list to store the key nodes and corresponding operation orders in all offset paths, and form a complete offset link data structure.
[0044] The specific steps of S4 are as follows: S401: Invoke the path offset trigger linked list, extract consecutive operation nodes, obtain the corresponding audio-visual content types, screen the content data with synchronizable characteristics, and obtain the audio-visual type matching set; First, traverse the operation sequence stored in the trigger linked list, identify the media content involved in each operation, determine whether it belongs to an audio or video element, and record the file format, duration, and playback status of these elements. For content of the audio-visual type, further screen whether it has synchronizable characteristics, that is, determine whether the audio and video can be time-axis aligned. If the start time of an audio file has a fixed offset from the corresponding video file, such as within 0.5 seconds, then it is considered to have synchronizable characteristics. If the offset time exceeds 2 seconds, then its time relationship needs to be further analyzed. Subsequently, screen the content data that meets the synchronizable characteristics, that is, extract all audio-visual content that can be synchronously played, and establish an audio-visual matching index. Finally, obtain the audio-visual type matching set, which stores all audio-visual elements that meet the synchronizable characteristics and their associated operation nodes.
[0045] S402: Based on the audio-visual type matching set, calculate the rendering delay and operation response time difference, extract the task loading timeliness data, calculate the synchronization difference between task loading and response, screen the task nodes that meet the synchronizable characteristics, and obtain the calculation result of the task synchronization difference; First, record the actual loading time of the audio - video file during playback, and measure the time interval from when the user triggers an operation to when the audio - video content rendering is completed. If the loading time of the audio - video is less than 500 milliseconds, and the time interval from user operation to rendering is between 300 milliseconds and 1 second, then it is determined that the audio - video responds quickly. If the loading time exceeds 1 second, and the operation response time interval exceeds 2 seconds, then it is judged to respond slowly. Then, extract the timeliness data of task loading, that is, traverse all audio - video playback tasks, count their loading durations, and calculate the synchronization difference between loading and response of the tasks, that is, the interval between when the user operation triggers playback and the actual playback time. If this interval is less than 1 second, then the task synchronization is considered good. If the interval exceeds 2 seconds, then timing compensation is required. After calculating the synchronization differences of all tasks, filter the task nodes that meet the synchronization characteristics, that is, find the tasks with synchronization errors within an acceptable range. Finally, obtain the calculation result of the task synchronization difference, where the synchronization error data of all audio - video tasks and the list of synchronizable tasks are stored.
[0046] S403: Call the calculation result of the task synchronization difference, judge the compliance of the channel multiplexing tolerance interval, extract the task indices that meet the multiplexing conditions, and filter the reusable task data to obtain the set of reusable task indices; First, determine the time tolerance standard for channel multiplexing, that is, set the maximum tolerance value within the allowable range of task synchronization error. If the synchronization error of a certain task is between 0 and 500 milliseconds, then it is determined that it can be multiplexed through the channel. If it exceeds 1 second, then it needs to be processed separately. Subsequently, extract the task indices that meet the multiplexing conditions, that is, filter out all tasks with synchronization errors within the set range and record their task numbers. Then, filter the reusable task data, that is, obtain all audio - video tasks that meet the tolerance standard and analyze their multiplexing priorities. If the synchronization error of a certain task is less than 200 milliseconds and belongs to the same playback group, then it is preferentially multiplexed. If the synchronization error of a certain task is around 500 milliseconds but has the same start time as other tasks, then it can still be multiplexed. Finally, obtain the set of reusable task indices, where the task numbers of all tasks that can be multiplexed through the channel and their corresponding multiplexing parameters are stored.
[0047] The specific steps of S5 are as follows: S501: Call the set of reusable task indices, extract the graphic - text combination information, identify the positions of the corresponding elements in the editing interface, and filter all the matched combined elements to obtain the graphic - text combination positioning data; First, traverse all task data stored in the task index set, filter tasks involving a combination of graphics and text, extract their associated image and text elements, obtain the coordinate information of each graphics and text element, and store the position information of each element in the current editing interface. Subsequently, compare with the original graphics and text structure diagram to match the relative positions of the graphics and text elements. If the distance between a certain text block and the corresponding image remains within the range of 20 to 50 pixels, it is judged as a valid match; if it exceeds 100 pixels, it is marked as a potentially misaligned element. Then, filter all the matched combined elements, that is, eliminate the graphics and text combinations with no association or large matching deviations. Finally, establish the graphics and text combination positioning data, which stores all the matched graphics and text combinations and their coordinate information in the current editing interface.
[0048] S502: Based on the graphics and text combination positioning data, calculate the combined distance between the current element position and the original order, filter the text blocks that meet the logical distance requirements, extract the recombined sequence of the text blocks, call the element order matched by the combined distance data, and obtain the text combination reconstruction sequence; First, obtain the current coordinates of all text blocks and calculate the displacement between them and the expected coordinates in the original order. If the displacement of a certain text block is between 10 and 30 pixels, it is judged that it basically maintains the original order; if it exceeds 50 pixels, further analyze its offset trend. Subsequently, filter the text blocks that meet the logical distance requirements, that is, eliminate the text blocks with too large displacements, and only retain the text blocks with small offsets and maintaining the relative order. Then, extract the recombined sequence of the text blocks, rearrange them according to the actual position order of the text blocks, and record the adjusted text order. If the order of a certain text block changes but the offset is within a reasonable range, add it to the recombined sequence. Finally, call the combined distance data, match the adjusted text element order, and store the text combination reconstruction sequence, which contains all text blocks that meet the logical order and their arrangement results.
[0049] S503: Call the text combination reconstruction sequence, re - establish the associated logical path, filter the adjusted graphics and text combination structure, extract the current document element sorting, and obtain the multimedia document structure rearrangement result; First, obtain the association information between all text blocks and images, and adjust the graphics and text combination structure according to the reconstruction sequence. If the relative position offset between a certain text block and the original image is within 30 pixels, retain the original association; if it exceeds 50 pixels, re - assign its corresponding image block. Subsequently, filter the adjusted graphics and text combination structure, that is, eliminate the text - image combinations with unreasonable associations due to position changes, and rearrange all valid combinations. Then, extract the sorting information of the current document elements, that is, obtain the final arrangement order of all texts and images, and store it in the multimedia document structure data. Finally, obtain the multimedia document structure rearrangement result, which stores all the adjusted graphics and text combination orders and their final positions in the document.
[0050] Please refer to Figure 2 , a multimedia document online editing system, comprising: The graphic and text structure analysis module obtains the pixel coordinates of the image boundary and the starting position of the corresponding text characters, calculates the coordinate difference in the alignment direction of the graphic and text blocks, generates a path sequence based on the offset direction of adjacent elements, matches the graphic and text combination tags and the editing attribute status identifier, identifies the graphic and text combination relationship, and obtains the initial graphic and text structure diagram; The view offset detection module calls the initial graphic and text structure diagram, detects the coordinate positions of the graphic and text elements in the view, compares with the original path vector, calculates the zoom ratio, scroll displacement and offset angle, judges the change status of the combination relationship, and obtains the path offset identifier set; The editing behavior analysis module calls the path offset identifier set, extracts the graphic and text combination index, collects the timestamp, layer level and action label of the user's editing operation, calculates the trigger interval and label fitting degree of adjacent actions, constructs a continuous operation node link, and generates a path offset trigger linked list; The task synchronization calculation module calls the path offset trigger linked list, analyzes the operations of the audio and video content, calculates the rendering delay and the operation response time difference, compares the synchronization difference between the loading time efficiency and the response rhythm, judges the channel multiplexing tolerance interval, and generates a reusable task index set; The document structure rearrangement module calls the reusable task index set, extracts the positions of the graphic and text elements in the editing interface, calculates the combination distance index from the original order, filters the text blocks with matching logical distances, reconstructs the graphic and text combination order and the associated logical path, and generates the multimedia document structure rearrangement result.
[0051] As described above, only the specific implementation manners of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A multimedia document online editing method, characterized in that: The following steps are involved: S1: Obtain the image boundary coordinates and the text starting position, extract the difference based on the alignment direction of the image and text, generate a path sequence based on the offset of adjacent elements, match the combined label and the editing attribute status, and construct the initial structure diagram of the image and text; S2: calling the initial structure graph path, comparing the current graph coordinates with the original path vector, calculating the offset angle and path difference through the zoom ratio and the rolling displacement, and generating a path offset identification set; S3: calling the image-text combination index in the offset identification center, collecting the timestamp, layer level and action label in the user editing, calculating the trigger spacing and label fitting degree between actions, screening continuous operation nodes, extracting the trigger sequence in the image-text combination, and generating a path offset trigger chain list; S4: calling the audio and video content type associated with the continuous operation node in the path offset trigger linked list, obtaining the rendering delay and response time difference, calculating the task synchronization deviation, determining whether the channel multiplexing tolerance interval is met, and generating a reusable task index set; S5: According to the reusable task index set, locate the corresponding elements of the editing interface, compare the combination distance between them and the original order, reconstruct the logical order of the text blocks, and generate a multimedia document structure rearrangement result.
2. The method for online editing of multimedia documents according to claim 1, characterized in that: The initial structure diagram of the graphic includes image boundary coordinates, text character positions, path direction information, combination tag types, and editing attribute status; the path offset identification set includes offset angle values, path distance difference values, scaling parameters, and scrolling displacements; the path offset trigger linked list includes operation time nodes, layer level identifications, action type tags, and trigger sequence information; the reusable task index set includes audio and video content types, rendering delay parameters, response time difference indicators, and channel multiplexing tolerance ranges; the multimedia document structure rearrangement results include element combination distances, text block sequence relationships, associated logical paths, and interface positioning positions.
3. The method for online editing of multimedia documents according to claim 1, characterized in that: The specific steps of S1 are: S101: Obtaining the pixel coordinates of the image boundary, matching the starting position of the text character, calculating the offset between the two, extracting the coordinate difference in the horizontal and vertical directions, establishing a coordinate matching relationship, and obtaining the pixel text matching offset; S102: based on the pixel text matching offset, extract the offset direction of adjacent characters, calculate the displacement difference, filter the data that meets the text block alignment direction, generate a character connection path, call the offset data to sort the path, and obtain a character offset path sequence; S103: calling the character offset path sequence, matching the image-text combination label with the edit attribute status identifier, analyzing the path change trend, screening the combination method that meets the structure, adjusting the image-text association relationship, and obtaining the image-text initial structure diagram.
4. The method for online editing of multimedia documents according to claim 1, characterized in that: The specific steps of S2 are: S201: calling the graphic-text combination path in the graphic-text initial structure diagram, obtaining the coordinate position of the graphic-text element of the current view, calculating the difference between the current view coordinate and the original path vector, establishing a corresponding relationship between the two, and obtaining graphic-text coordinate comparison data; S202: Based on the image-text coordinate comparison data, calculate the overall zoom and scroll offset index, extract the offset angle and path distance difference of adjacent data points in the coordinate transformation process, analyze the path change trend, filter the change data of the path offset direction, and obtain the path offset calculation result; S203: calling the path offset calculation result, analyzing the structural change of the combined path, identifying the adjustment state of the graphic-text combination relationship, screening the feature points of the changed path, extracting the corresponding offset mode, and obtaining a path offset identification set.
5. The method for online editing of multimedia documents according to claim 4, characterized in that: The overall zoom scroll offset index calculation formula is specifically: ; in, Represents the overall zoom scroll offset indicator, Represents the current view The horizontal coordinate value of the coordinate point, Represents the original view The horizontal coordinate value of the coordinate point, Represents the current view The vertical coordinate value of the coordinate point, Represents the original view The vertical coordinate value of the coordinate point, Represents the original view The path distance of the coordinate points, To prevent small positive numbers with zero denominators, Represents the total number of coordinate data points, Represents the current view Roll displacement values, Represents the original view Roll displacement values, Represents the total number of rolling displacement data points.
6. The method for online editing of multimedia documents according to claim 1, characterized in that: The specific steps of S3 are: S301: calling the path offset identification set, extracting the image-text combination index with path offset, screening the corresponding image-text combination data, establishing an index list, and obtaining the image-text offset index set; S302: Based on the image-text offset index set, collect the timestamp, layer level, and action label in the user editing operation, calculate the average trigger time interval, filter the action sequence that conforms to the path offset rule, analyze the fit between the labels, identify the association relationship between the actions, and obtain the action trigger matching data; S303: calling the action trigger matching data, constructing a continuous operation node link, extracting the trigger sequence in the graphic combination, screening the key nodes of the path deviation, and obtaining the path deviation trigger chain list.
7. The method for online editing of multimedia documents according to claim 6, characterized in that: The calculation formula for the average trigger time interval is specifically: ; in, Representative The average triggering time interval of the operation relative to the previous operation, Represents the current state The timestamp of the operation. Representative The timestamp of the operation. Represents the adjustment factor of layer level changes, Represents the current state The layer level of the layer, Representative The layer level of the layer, Represents the number of operations recorded in the interface. Represents the current state The trigger path offset of the operation, Represents the original state The trigger path offset of the operation, Represents the total number of path offset data points involved in the interface.
8. The method for online editing of multimedia documents according to claim 1, characterized in that: The specific steps of S4 are: S401: calling the path offset trigger linked list, extracting continuous operation nodes, obtaining corresponding audio and video content types, screening content data with synchronizable characteristics, and obtaining an audio and video type matching set; S402: Based on the audio and video type matching set, calculate the rendering delay and the operation response time difference, extract the task loading time efficiency data, calculate the synchronization difference between the task loading and the response, select the task nodes that meet the synchronization characteristics, and obtain the task synchronization difference calculation result; S403: calling the task synchronization difference calculation result, determining compliance with the channel reuse tolerance interval, extracting task indexes that meet the reuse conditions, screening reusable task data, and obtaining a reusable task index set.
9. The method for online editing of multimedia documents according to claim 1, characterized in that: The specific steps of S5 are: S501: calling the reusable task index set, extracting the image-text combination information, identifying the corresponding element position in the editing interface, screening all matched combination elements, and obtaining the image-text combination positioning data; S502: Based on the image-text combination positioning data, calculate the combination distance between the current element position and the original order, select the text blocks that meet the logical distance requirements, extract the reorganized sequence of the text blocks, call the element sequence matched by the combination distance data, and obtain the text combination reconstruction sequence; S503: calling the text combination reconstruction sequence, re-establishing the associated logical path, screening the adjusted graphic and text combination structure, extracting the current document element sequence, and obtaining the multimedia document structure rearrangement result.
10. A multimedia document online editing system, characterized in that: According to a multimedia document online editing method according to any one of claims 1 to 9, the system comprises: The image-text structure parsing module obtains the image boundary pixel coordinates and the corresponding text character starting position, calculates the coordinate difference of the image-text block alignment direction, generates a path sequence according to the offset direction of adjacent elements, matches the image-text combination label and the edit attribute status identifier, identifies the image-text combination relationship, and obtains the image-text initial structure diagram; The view offset detection module calls the image and text initial structure diagram, detects the coordinate position of the image and text elements in the view, compares it with the original path vector, calculates the zoom ratio, scroll displacement and offset angle, determines the change state of the combination relationship, and obtains the path offset identification set; The editing behavior analysis module calls the path offset identification set, extracts the image-text combination index, collects the timestamp, image level and action label of the user's editing operation, calculates the trigger spacing and label fitting degree of adjacent actions, constructs a continuous operation node link, and generates a path offset trigger chain list; The task synchronization calculation module calls the path offset trigger chain list, analyzes the operation of the audio and video content, calculates the rendering delay and the operation response time difference, compares the synchronization difference between the loading time efficiency and the response rhythm, determines the channel reuse tolerance interval, and generates a reusable task index set; The document structure rearrangement module calls the reusable task index set, extracts the position of the graphic and text elements in the editing interface, calculates the combination distance index with the original order, screens the text blocks with matching logical distances, reconstructs the graphic and text combination order and the associated logical path, and generates the multimedia document structure rearrangement result.
Citation Information
Patent Citations
Multimedia data processing method, device and equipment and readable storage medium
CN114257843A
Document editing method and device, medium and electronic equipment
CN117010336A
Face recognition application method and system combined with intelligent glasses
CN119828856A
System and management of semantic indicators during document presentations
US20200159838A1
Cited By
Cloud platform auditing method based on collaborative auditing
CN121073058A