A method, equipment, media, and program product for on-the-fly typesetting.

By transforming images into structured data and dividing them into main viewing areas, secondary viewing areas, and avoidance areas in complex application scenarios, and constructing semantic adhesion clusters by combining spatial adjacency, the problem of image and text elements being fragmented or falling into visual blind spots during 3D pasting is solved, thus achieving the integrity and accuracy of image and text information.

CN122018830BActive Publication Date: 2026-07-17BEIJING SHUOFANG INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING SHUOFANG INFORMATION TECH CO LTD
Filing Date
2026-04-10
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In complex application scenarios, existing image recognition and typesetting methods cannot effectively prevent text and image elements from being cut off by fold lines or falling into visual blind spots during the 3D pasting process, resulting in a decrease in the accuracy of information transmission.

Method used

By converting the original captured images into structured data, predicting the physical deformation boundary using the target's 3D topological morphology, and dividing it into a main view area, a secondary view area, and an avoidance area, semantic adhesion clusters are constructed by combining spatial adjacency, and local scaling and offset adjustments are performed to ensure the integrity and accuracy of the image and text module in complex 3D pasting scenes.

Benefits of technology

It effectively avoids the visual blind spots and semantic fragmentation problems of traditional two-dimensional typesetting in complex three-dimensional pasting scenarios, and improves the integrity and accuracy of graphic and textual information transmission of special labels in actual physical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122018830B_ABST
    Figure CN122018830B_ABST
Patent Text Reader

Abstract

A method, equipment, media, and program product for on-demand image printing layout are disclosed, relating to the fields of image processing and intelligent printing technology. The layout equipment, based on converting the original image into structured data and performing global mapping, divides the consumable into a main viewing area, a secondary viewing area, and an avoidance area. It also constructs semantically connected clusters by combining the spatial adjacency of objects. This allows it to detect whether logically related graphic modules will be interrupted by fold lines or obscured by holes. If a conflict occurs, it breaks the global scaling constraint and independently calculates and adjusts the local scaling ratio and offset coordinates for that cluster. This derivation process of multi-dimensional feature interaction effectively avoids the visual blind spots and semantic fragmentation problems of traditional two-dimensional layout in complex three-dimensional pasting scenarios, improving the integrity and accuracy of graphic information transmission for special labels (such as flag-shaped and wrapped labels) in practical physical applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of image processing and intelligent printing technology, and in particular to a method, typesetting equipment, media, and program products for instant typesetting by shooting. Background Technology

[0002] With the development of mobile internet and portable printing devices, mobile photo-and-print technology has been widely used in asset management, warehousing and logistics, and data center cable labeling. Users often need to use applications installed on smartphones and other mobile devices to photograph old equipment nameplates, manuals, or paper labels, and then quickly copy and print the extracted text and images onto new, dedicated blank label supplies.

[0003] In related technologies, image recognition algorithms are often used to extract various text and graphic elements from the captured image, generating their respective independent two-dimensional bounding rectangles. Then, the absolute physical width and height of the consumables loaded within the currently connected printer are obtained. The global scaling ratio between the overall size of the source image and the plane size of the target consumables is calculated. Based on this scaling ratio and two-dimensional geometric collision rules, all independent text and graphic elements are treated as free layout blocks, sequentially scaled, and tiled to fill the two-dimensional printable plane of the target consumables. This process generates the final printed image data while ensuring that the text and graphic elements do not exceed the physical boundaries of the consumables.

[0004] However, in complex application scenarios such as power line inspections or communication equipment rooms, printed special labels (such as flag-shaped labels, wrap-around labels, or folded labels) are often not pasted flat, but require physical deformation steps such as folding or wrapping to be attached to the area. The geometric scaling and flow avoidance of related technologies may cause related graphic elements that are originally closely connected in the source image (such as warning triangle icons and their bottom explanatory text) to be unintentionally assigned to either side of the label's fold line or curved extension line during the layout adaptation process, reducing the accuracy of information transmission. Summary of the Invention

[0005] This application provides a method, equipment, media, and program product for on-the-fly typesetting of images and text, which can improve the accuracy of information transmission of image and text information in complex application scenarios.

[0006] Firstly, this application provides a method for on-demand typesetting based on a photograph, applied to a typesetting device. The method includes: inputting an original photographed image into a preset image recognition model, outputting a structured data list containing multiple preset object types, wherein each object in the structured data list contains corresponding source image data, a corresponding type identifier, and mapping space coordinate parameters in the original photographed image; determining the global scaling factor of the original photographed image in a preset two-dimensional target printing coordinate system based on the document boundary, and then using the global scaling factor to map the mapping space coordinate parameters of each object in the structured data list to the two-dimensional target printing coordinate system to generate preliminary layout data, wherein the document boundary is determined based on the object distribution characteristics in the structured data list; and determining the printing consumption based on the original photographed image and preset application scenario configuration instructions. The physical deformation boundary of the material under the target three-dimensional topology is determined, and based on this physical deformation boundary, the printable area of ​​the material to be printed is divided into a main view area and a secondary view area at non-coplanar viewing angles, as well as an avoidance area where there are structural obstructions. The mapped spatial coordinate parameters of each object in the structured data list are traversed, the spatial adjacency between two objects is calculated, and multiple objects with spatial adjacency higher than a preset threshold are identified as semantically connected clusters. If it is determined that the projection area of ​​a target semantically connected cluster after being mapped by the global scaling factor crosses the boundary between the main view area and the secondary view area or covers the avoidance area, the local scaling ratio and offset coordinates of the target semantically connected cluster falling into the main view area or the secondary view area are calculated. The preliminary layout data is updated according to the local scaling ratio and the offset coordinates to obtain the final layout data.

[0007] By adopting the above technical solution, the typesetting device, after converting the original image into structured data and performing global mapping, divides the consumables into a main viewing area, a secondary viewing area, and an avoidance area, and constructs semantically connected clusters based on the spatial adjacency between objects. This allows it to detect whether logically related graphic modules will be interrupted by fold lines or obscured by holes. If a conflict occurs, the global scaling limitation is broken, and the local scaling ratio and offset coordinates are independently calculated and adjusted for that connected cluster. This derivation process of multi-dimensional feature interaction effectively avoids the visual blind spots and semantic fragmentation problems of traditional two-dimensional typesetting in complex three-dimensional pasting scenarios, improving the integrity and accuracy of graphic information transmission for special labels (such as flag-shaped and wrapped labels) in practical physical applications.

[0008] In conjunction with some embodiments of the first aspect, in some embodiments, the step of determining the physical deformation boundary of the consumable to be printed under the target three-dimensional topological shape based on the original captured image and the preset application scenario configuration instructions, and dividing the printable area of ​​the consumable to be printed into a main viewing area and a secondary viewing area at non-coplanar viewing angles, as well as an avoidance area with structural obstructions based on the physical deformation boundary, specifically includes: determining the target three-dimensional topological shape of the target physical carrier corresponding to the document and estimating the cross-sectional size based on the background environment image features located outside the document boundary in the original captured image; determining a virtual three-dimensional topological shape matching the target three-dimensional topological shape from a preset consumable folding mapping library. The virtual three-dimensional folding model is simulated by substituting the estimated cross-sectional dimensions and the two-dimensional geometric dimensions of the consumable to be printed into the virtual three-dimensional folding model. The projection coordinates of the virtual three-dimensional folding model on the two-dimensional unfolded plane are determined and the projection coordinates are used as the physical deformation boundary of the consumable to be printed under the target topological shape. Using the physical deformation boundary as the dividing line, the continuous area in the printable area of ​​the consumable to be printed that is located on the starting plane and has a single orthogonal viewing angle is divided into the main viewing area. The continuous area that is deflected at an angle due to crossing the physical deformation boundary is divided into the auxiliary viewing area. The reserved assembly hole area and physical gap area on the consumable to be printed are divided into the avoidance area.

[0009] By employing the above technical solution, the typesetting equipment utilizes the background environmental image features in the original captured images to deduce the three-dimensional topological shape and cross-sectional dimensions of the target physical carrier. These physical parameters are then substituted into a virtual three-dimensional folding model for geometric simulation, calculating the physical deformation boundaries of the consumable during actual bending. Based on this, the two-dimensional plane is scientifically reduced in dimension to divide into a main viewing area, a secondary viewing area, and an avoidance area. This allows the typesetting algorithm to proactively avoid visual blind spots and physically damaged areas, improving the scientific rigor and rational space utilization of complex carrier label typesetting.

[0010] In conjunction with some embodiments of the first aspect, in some embodiments, the step of determining the target three-dimensional topological shape and estimating the cross-sectional size of the target physical carrier corresponding to the document based on the background environment image features located outside the document boundary in the original captured image specifically includes: extracting the background environment image features located outside the document boundary in the original captured image, and performing edge gradient detection on the grayscale matrix corresponding to the background environment image features to determine an environment edge binary map; performing a Hough line transform on the environment edge binary map with a preset angle resolution and a preset distance resolution to determine multiple sets of parametric lines with accumulator peak values ​​greater than a preset truncation threshold; and determining from the multiple sets of parametric lines a parallelism error lower than an allowable threshold and exceeding the limit. The document boundary is defined by parallel contour boundary segments. Based on the contour distortion rate of the closed region enclosed by the parallel contour boundary segments, the closed region is mapped to a preset geometric primitive library for matching to determine the target's three-dimensional topological shape. After obtaining the reference physical size carried in the preset application scenario configuration instruction, the mapping ratio coefficient between the reference physical size and the first pixel span value of the document boundary in the original captured image is calculated. The reference physical size is the known one-dimensional absolute length of the document in the real world. Based on the second pixel span value of the parallel contour boundary segments in the original captured image, the product of the mapping ratio coefficient and the second pixel span value is determined as the estimated cross-sectional size.

[0011] By employing the above technical solutions, the typesetting equipment, through edge gradient detection and Hough linear transformation, can extract the parallel contour boundary segments of the target carrier from complex background interference. Combining contour deformation distortion rate with feature matching from a pre-set geometric primitive library, the three-dimensional topological morphology of the carrier is locked, improving the scientific rigor of complex carrier label typesetting.

[0012] In conjunction with some embodiments of the first aspect, in some embodiments, the step of mapping the closed region to a preset geometric primitive library for matching based on the contour distortion rate of the closed region enclosed by the parallel contour boundary line segment to determine the target three-dimensional topological shape specifically includes: after drawing orthogonal projection rays from the vertices of the document boundary to the parallel contour boundary line segment, if it is determined that the topological closed region enclosed by the orthogonal projection rays and the parallel contour boundary line segment includes the document boundary, the topological closed region is taken as the two-dimensional orthogonal projection region of the target physical carrier; extracting the light and shadow gradient vectors of the background pixels in the two-dimensional orthogonal projection region, and combining them with the contour distortion rate of the topological closed region to form a region structure descriptor; performing feature space matching between the region structure descriptor and the preset geometric primitive library to determine the geometric solid with the highest confidence interval as the target three-dimensional topological shape.

[0013] By adopting the above technical solution, the typesetting device integrates the light and shadow features reflecting surface curvature with the contour distortion rate reflecting perspective relationships to construct a multi-dimensional region structure descriptor, and performs high-confidence feature space matching in a preset geometric primitive library. This derivation process, which combines geometric topology with optical rendering features, reduces the risk of failure of single edge detection in low-light, reflective, or complex texture environments.

[0014] In conjunction with some embodiments of the first aspect, in some embodiments, the step of calculating the local scaling ratio and offset coordinates of the target semantic cluster falling into the main view area or the secondary view area specifically includes: based on the global bounding rectangle of the target semantic cluster in the preliminary layout data, extracting the initial cluster width, initial cluster height, and initial center point coordinates from the global bounding rectangle; determining a target area larger than the global bounding rectangle, and determining the effective area width, effective area height, and target center point coordinates of the target area, wherein the target area is the unoccupied main view area or the secondary view area; while maintaining the initial cluster width and the initial... Under the constraint of a constant cluster height ratio, calculate a first ratio of the effective area width to the initial cluster width, and a second ratio of the effective area height to the initial cluster height; determine the local scaling ratio based on the smaller of the first and second ratios and a preset safety margin coefficient; after scaling the global bounding rectangle according to the local scaling ratio to obtain the target bounding rectangle, calculate the required horizontal and vertical translation distances when aligning the center point of the target bounding rectangle to the coordinates of the target center point, and determine the two-dimensional vector containing the horizontal and vertical translation distances as the offset coordinates.

[0015] By employing the above technical solution, when the target semantically connected clusters face the risk of cross-regional fragmentation, the typesetting device extracts the geometric parameters of its global bounding rectangle and the target safe area. Under the constraint of maintaining the initial aspect ratio, it calculates the smaller of the bidirectional ratio between the effective area and the initial cluster, and combines this with a safety margin coefficient to determine the local scaling ratio, thereby deriving the translation offset coordinates. This process of feature interaction ensures that when the graphic module undergoes spatial migration and size compression, there will be no visual distortion of the aspect ratio, nor will there be printing overflow due to being too close to the boundary. This allows highly related graphic information to be completely embedded in a single visible area at a reasonable proportion, improving the local typesetting quality of the label.

[0016] In conjunction with some embodiments of the first aspect, in some embodiments, the step of traversing the mapped spatial coordinate parameters of each object in the structured data list, calculating the spatial adjacency between two objects, and identifying multiple objects with spatial adjacency higher than a preset threshold as a semantically cohesive cluster specifically includes: traversing the mapped spatial coordinate parameters of each object in the structured data list and calculating the spatial physical distance between two objects; determining the corresponding semantic association coefficient from a preset type association matrix based on the type identifiers of the two objects, wherein the elements in the type association matrix define the logical dependence degree between different preset object types; calculating the spatial physical distance by weighting the spatial physical distance using the semantic association coefficient to obtain the spatial adjacency between the two objects; constructing an object topology graph by establishing undirected edges between two nodes with spatial adjacency higher than the preset threshold, using all objects in the structured data list as nodes; extracting each maximum connected subgraph from the object topology graph and identifying each maximum connected subgraph as the semantically cohesive cluster.

[0017] By adopting the above technical solution, the typesetting device introduces a pre-set type association matrix. The system extracts the type identifier of objects to obtain semantic association coefficients, and then weights and fuses these coefficients with absolute spatial physical distance to calculate a comprehensive spatial adjacency. Based on this adjacency, an object topology graph is constructed and the maximum connected subgraph is extracted, achieving intelligent clustering of text and image elements. This reduces semantic breaks and information misleading caused by typesetting reorganization.

[0018] In conjunction with some embodiments of the first aspect, in some embodiments, the step of updating the preliminary layout data according to the local scaling ratio and the offset coordinates to obtain the final layout data specifically includes: performing spatial adjustment on the objects contained in the target semantic adhesion cluster in the preliminary layout data based on the local scaling ratio and the offset coordinates to generate an updated cluster placeholder bounding box region; traversing discrete objects in the preliminary layout data that do not belong to the target semantic adhesion cluster, performing a Boolean intersection operation between the mapped spatial coordinate parameters of the discrete objects and the updated cluster placeholder bounding box region to determine the affected discrete objects with overlapping positions; using the updated cluster placeholder bounding box region and the avoidance area as a joint obstacle boundary, determining the blank projection point with the shortest distance to the mapped spatial coordinate parameters of the affected discrete object; generating secondary offset coordinates to move the affected discrete object to the corresponding blank projection point, and using the secondary offset coordinates to correct the position of the affected discrete object in the preliminary layout data; and recombining the spatially adjusted data of the target semantic adhesion cluster with the data of all discrete object data after position correction to obtain the final layout data.

[0019] By adopting the above technical solution, after completing the local adjustment of the target semantic clusters, the typesetting equipment merges the updated cluster placeholder outline with the inherent avoidance area into a joint obstacle boundary, and finds blank projection points for the affected objects based on the shortest distance principle, generating secondary offset coordinates for linkage correction. While ensuring the integrity of the core semantic module, it improves the density of graphic and textual information layout within the limited printing space.

[0020] In a second aspect, this application provides a typesetting device comprising: one or more processors and a memory; the memory is coupled to the one or more processors and is used to store computer program code, the computer program code including computer instructions, wherein the one or more processors invoke the computer instructions to cause the typesetting device to perform the method described in the first aspect and any possible implementation thereof.

[0021] Thirdly, this application provides a computer program product containing instructions that, when run on a typesetting device, cause the typesetting device to perform the method described in the first aspect and any possible implementation thereof.

[0022] Fourthly, this application provides a computer-readable storage medium including instructions that, when executed on a typesetting device, cause the typesetting device to perform the method described in the first aspect and any possible implementation thereof.

[0023] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0024] 1. By employing a technique that predicts physical deformation boundaries based on the target's 3D topological morphology and divides the visible area, while simultaneously using spatial adjacency to construct semantically linked clusters for local scaling and offset adjustments, the typesetting device can anticipate the spatial folding state of consumables during actual pasting. When it detects logically related graphic modules (i.e., semantically linked clusters) crossing the visible area boundary or covering avoidance zones, it breaks global scaling limitations and adaptively adjusts the local scaling ratio and offset coordinates. This process effectively solves the technical problem in existing 2D typesetting technologies where relying solely on geometric collision avoidance in complex 3D pasting scenarios easily leads to related graphics being severed by fold lines or falling into visual blind spots. This ensures the integrity and readability of graphic information for special labels in actual physical applications, improving the accuracy of information transmission in complex scenarios.

[0025] 2. By employing techniques that utilize background image features to deduce the three-dimensional topology and cross-sectional dimensions of the target physical carrier, and then using a virtual three-dimensional folding model to calculate two-dimensional projected coordinates as physical deformation boundaries for region division, the typesetting equipment can transform background information from the original image into specific physical parameters. Through geometric simulation, it can accurately calculate the crease positions generated when the consumable is actually bent, and accordingly, reduce the dimensionality of the two-dimensional plane to a main viewing area, a secondary viewing area, and an avoidance area, transforming the real-world three-dimensional installation constraints into two-dimensional typesetting space rules. This achieves the technical effect of providing clear spatial constraint boundaries for graphic and text typesetting, proactively avoiding visual blind spots and physically damaged areas, and improving the rational utilization of space in complex carrier label typesetting.

[0026] 3. By adopting a technique that determines the local scaling ratio based on the minimum bidirectional size ratio of the target region and the semantically connected cluster and the safety margin coefficient while maintaining the initial aspect ratio, and then calculates the center alignment offset coordinates accordingly, the local layout quality of the label is improved. Attached Figure Description

[0027] Figure 1 This is a flowchart illustrating a method for on-the-spot typesetting in an embodiment of this application;

[0028] Figure 2 This is another flowchart illustrating a method for typesetting and printing on the spot, as described in this application embodiment;

[0029] Figure 3 This is a schematic diagram of an exemplary hardware structure of the typesetting device in the embodiments of this application. Detailed Implementation

[0030] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to and includes any or all possible combinations of one or more of the listed items.

[0031] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.

[0032] Please see Figure 1This is a flowchart illustrating a method for typesetting by shooting and printing in an embodiment of this application.

[0033] S101. Input the original captured image into the preset image recognition model and output a structured data list containing multiple preset object types.

[0034] The original captured image refers to the initial two-dimensional pixel matrix containing the text and image information to be copied, acquired by the typesetting equipment through the image acquisition module. It represents the visual image of an old equipment nameplate or paper label taken by the user on-site. The preset image recognition model refers to a deep learning network architecture pre-trained on a large amount of labeled image data. It represents a set of algorithms capable of extracting features from complex backgrounds and performing classification and localization, such as a convolutional neural network-based object detection model. The preset object type refers to a pre-defined set of image and text element classification labels, representing different semantic entities that may exist in the image, such as text lines, barcodes, QR codes, warning icons, or company logos. The structured data list refers to a set of unstructured image pixels converted into a computer-parseable data set with fixed field formats, representing the ordered arrangement of the recognition results. The source image data refers to the local pixel content corresponding to a specific object, cropped or extracted from the original captured image, representing the visual appearance of the object itself. The type identifier refers to the classification attribute code assigned to each recognized object, indicating which preset object type the object belongs to. Mapped space coordinate parameters refer to the position and size information of the identified object in the two-dimensional pixel coordinate system of the original captured image. They are used to represent the geometric attributes of the object's bounding box, such as the coordinates of the top left corner, width, and height.

[0035] In scenarios where users take photos of old labels in complex environments using the camera of a mobile terminal or typesetting equipment, triggering a "snap and print" task, the typesetting equipment first receives the raw image from the image acquisition hardware and loads it into memory for preprocessing. This includes scaling the image to the standard input resolution required by the preset image recognition model and normalizing pixel values ​​to eliminate the effects of uneven lighting. Subsequently, the typesetting equipment inputs the preprocessed image tensor into the forward propagation network of the preset image recognition model. The model's underlying feature extraction backbone network progressively extracts edge, texture, and high-level semantic features of the image through multiple convolutional and pooling operations, generating multi-scale feature maps. Based on these feature maps, the model predicts bounding boxes of potential objects at different spatial locations and scales, and outputs a confidence score for each bounding box indicating that it contains a specific preset object type. The typesetting equipment uses a non-maximum suppression algorithm to filter out redundant bounding boxes with excessive overlap and low confidence, retaining the optimal detection results. For each retained bounding box, the typesetting device extracts its internal pixels as source image data, records its corresponding classification result as a type identifier, and maps the bounding box's position back in the original unscaled image to generate mapping space coordinate parameters. Finally, the typesetting device encapsulates these attribute information of all identified objects and appends them sequentially to a dynamic array in memory according to a preset data structure protocol, thereby outputting a structured data list containing multiple preset object types, providing basic data support for subsequent typesetting mapping.

[0036] In some embodiments, the process of inputting the original captured image into a preset image recognition model and outputting a structured data list can be implemented in various ways: Optionally, the typesetting device adopts a cascaded object detection architecture based on region proposal for processing. The typesetting device first generates a large number of candidate anchor boxes that may contain objects by sliding the region proposal network on the feature map, and performs preliminary binary classification and position fine-tuning on these anchor boxes. Then, it uses the region of interest pooling technique to map candidate regions of different sizes into feature vectors of fixed dimensions. Finally, it inputs these feature vectors into a fully connected layer for multi-class classification and bounding box regression, thereby accurately outputting the type identifier and mapping space coordinate parameters of each object, and decoding the extracted local feature map into source image data and storing it in the structured data list. Optionally, the typesetting device employs a visual transformer architecture based on a self-attention mechanism. The device segments the original captured image into multiple fixed-size image blocks and flattens them into a sequence. Input terms are generated through linear projection and positional encoding. Then, a multi-layer, multi-head self-attention encoder captures the global contextual dependencies between image blocks. Subsequently, a decoder, combined with learnable object query vectors, directly predicts a fixed number of unordered prediction results. Finally, a bipartite graph matching algorithm optimally assigns the prediction results to the real targets. After removing background predictions, the category probabilities and normalized coordinates of the effective prediction results are converted into type identifiers and mapping space coordinate parameters, thereby constructing a structured data list. It is understood that other methods can also be used to achieve image feature extraction and structured data generation; this is not limited here. It should be noted that when generating the structured data list, the typesetting device also performs binarization and morphological denoising on the source image data to ensure the clarity of the subsequent printed output. Especially for text-type objects, an optical character recognition engine is additionally invoked to extract the corresponding string text content, which is stored as a supplementary field to the source image data in the structured data list.

[0037] S102. After determining the global scaling factor of the original captured image in the preset two-dimensional target printing coordinate system based on the document boundary, the global scaling factor is used to map the mapping space coordinate parameters of each object in the structured data list to the two-dimensional target printing coordinate system to generate preliminary layout data.

[0038] In this context, the document boundary refers to a virtual geometric bounding box determined by the typesetting device through analysis of the distribution range of all objects identified in two-dimensional space. It represents the effective content area in the original captured image that actually contains valid text and image information. The preset two-dimensional target printing coordinate system is a mathematical coordinate system corresponding to the physical plane of the blank consumable currently loaded in the printing device. It represents the absolute physical space of the final printed output, with its origin typically located at the upper left corner of the consumable. The global scaling factor is the overall size mapping ratio between the effective content area in the source image and the printable area of ​​the target consumable on the two-dimensional plane. It represents the uniform magnification or reduction factor required to completely fit all identified content into the consumable. Preliminary layout data refers to the set of temporary positions and dimensions of each object in the target printing coordinate system after global coordinate transformation. It represents the idealized two-dimensional typesetting state before considering the three-dimensional physical deformation of the consumable. Object distribution characteristics refer to the spatial topological attributes such as clustering, dispersion, and extreme points exhibited by the mapped spatial coordinate parameters of each object in the structured data list on the two-dimensional plane. They represent the overall arrangement pattern of the text and image content.

[0039] The typesetting device first iterates through the mapping space coordinate parameters of all objects in the structured data list, extracting the coordinate values ​​of the left, right, top, and bottom boundaries of each object's bounding box. By comparing these coordinate values, the typesetting device finds the minimum and maximum x-coordinates, minimum and maximum y-coordinates among all objects, and uses these four extreme coordinates to construct a minimum regular rectangle that completely encloses all objects, defining this rectangle as the document boundary. Next, the typesetting device obtains the maximum available width and maximum available height of the preset two-dimensional target printing coordinate system, calculating the ratio of the target available width to the document boundary width, and the ratio of the target available height to the document boundary height. To ensure that the text and image content does not suffer aspect ratio distortion after mapping and does not exceed the printing boundary, the typesetting device selects the smaller of these two ratios as the global scaling factor. Subsequently, using the top-left corner of the document boundary as the origin, the typesetting device performs an affine transformation on the mapping space coordinate parameters of each object in the structured data list using the calculated global scaling factor. Specifically, the original horizontal and vertical coordinates of each object are subtracted from the corresponding reference coordinates of the document boundary, and then multiplied by a global scaling factor to obtain its new coordinates in the preset two-dimensional target printing coordinate system. Simultaneously, its original width and height are also multiplied by the global scaling factor to obtain new dimensions. The typesetting equipment summarizes these new coordinates and dimensions after uniform scaling and translation calculations to generate preliminary layout data. At this point, all graphic objects maintain the same relative position and proportion on the target consumable plane as in the original captured image.

[0040] In some embodiments, the process of determining the global scaling factor and generating preliminary layout data based on the document boundary can be implemented in several ways: Optionally, the typesetting device adopts a compact boundary determination and mapping logic based on the convex hull algorithm. The typesetting device extracts the vertex coordinate set of all objects, uses the Graham scan algorithm to calculate the smallest convex polygon containing all vertices as the document boundary, then calculates the length and width of the smallest bounding rectangle of the convex polygon, compares it with the physical length and width of the preset two-dimensional target printing coordinate system to obtain the global scaling factor, and then constructs a homogeneous coordinate transformation matrix by calculating the offset between the centroid of the convex polygon and the center of the target printing coordinate system, combined with the global scaling factor. The matrix is ​​then used to perform batch matrix multiplication operations on the coordinates of all objects to directly output preliminary layout data containing rotation and translation information. Optionally, the typesetting device employs a content density-based weighted boundary determination and mapping logic. The device uses a kernel density estimation algorithm to calculate the two-dimensional density distribution of the center point coordinates of all objects, identifying the core content area where the density peaks are located. Isolated edge objects with densities below a preset noise threshold are considered interference and discarded. The document boundary is determined only based on the extreme coordinates of objects within the core content area. Then, a global scaling factor is calculated based on the size of the target printing coordinate system. Subsequently, when mapping coordinates using the global scaling factor, a centripetal contraction algorithm based on a gravity model is applied to the discarded isolated edge objects, causing them to automatically converge towards the core content area when mapped to the target coordinate system. Finally, the coordinates of the core objects and the contracted edge objects are merged to generate preliminary layout data. It is understood that other methods can also be used to achieve coordinate system mapping and layout generation; this is not limited here. It should be noted that when calculating the global scaling factor, the typesetting device also pre-deducts a preset physical printing blind zone (such as the paper feed margin at the edge of the consumable) from the absolute size of the preset two-dimensional target printing coordinate system to ensure that no object in the generated preliminary layout data falls within the printer's physically unprintable area.

[0041] S103. Based on the original captured image and the preset application scenario configuration instructions, determine the physical deformation boundary of the consumable to be printed under the target three-dimensional topological shape, and divide the printable area of ​​the consumable to be printed into the main viewing area and the auxiliary viewing area at non-coplanar viewing angles, as well as the avoidance area where there is structural obstruction, based on the physical deformation boundary.

[0042] The preset application scenario configuration command refers to the set of control parameters for the special purpose of the current printing task, which are input by the user through the interactive interface or automatically obtained by the system through the identification of the consumable RFID tag. This indicates the physical environment in which the tag will be applied, such as "cable entanglement" or "pipe folding." The target 3D topology refers to the three-dimensional geometric shape of the consumable after it is actually pasted or installed on the target object. This represents the bending, folding, or wrapping state of the consumable in three-dimensional space, such as a cylindrical surface or a V-shaped fold. The physical deformation boundary refers to the projection line of the curvature abrupt change or folding edge in the target 3D topology onto the two-dimensional unfolded plane of the consumable. This represents the geometric boundary line where the consumable surface undergoes spatial angular deflection. The printable area of ​​the consumable refers to the total two-dimensional planar area of ​​the consumable surface where printing ink or ribbon can be effectively adhered. The main viewing area refers to the continuous two-dimensional area of ​​the consumable in its three-dimensional form, facing the observer, with minimal change in surface normal vectors, providing a complete and distortion-free visual experience. The secondary viewing area refers to the continuous two-dimensional region of the consumable in three-dimensional form that deviates from the main viewing angle due to crossing the physical deformation boundary, requiring the observer to change their viewing angle to see it clearly. The avoidance area refers to pre-existing physical holes, die-cut gaps, or areas on the consumable that will be obscured after three-dimensional installation and cannot be used to display information. Background environment image features refer to the pixel texture, edge, and lighting information in the original captured image, excluding the effective text and image content, used to represent the appearance attributes of the target physical carrier (such as cables or pipes). Estimated cross-sectional dimensions refer to the geometric width or diameter of the target physical carrier perpendicular to its extension direction, calculated from the image. A virtual three-dimensional folding model refers to a mathematical and geometric model constructed in computer memory that simulates the physical deformation process of the consumable.

[0043] After generating the initial layout data, the typesetting device first parses the preset application scenario configuration instructions to clarify the specific physical constraints of the current task. The device extracts background environment image features located outside the document boundary from the original captured image. By analyzing the texture direction, light and shadow gradations, and edge contours of these background pixels, it infers the three-dimensional shape of the target physical carrier to which the old label is attached, thereby determining the target three-dimensional topological form of the document's corresponding target physical carrier. Combined with the pixel ratios in the image, it calculates the estimated cross-sectional dimensions of the carrier. Next, the typesetting device uses its built-in physics engine or a lookup table method to retrieve a virtual three-dimensional folding model matching the target three-dimensional topological form from a preset consumable folding mapping library. The typesetting device uses the calculated estimated cross-sectional dimensions and the two-dimensional geometric dimensions of the consumable to be printed as input parameters, substituting them into the virtual three-dimensional folding model for geometric simulation. By simulating the process of consumables surrounding or attaching to the target carrier, the typesetting equipment can calculate where the consumables will bend or experience abrupt changes in curvature. It then reverses the unfolding of these creases or curved surfaces in three-dimensional space, determining their projection coordinates on a two-dimensional unfolded plane. These projection coordinates constitute the physical deformation boundaries of the consumables to be printed under the target topological shape. Subsequently, the typesetting equipment uses these one or more physical deformation boundaries as dividing lines to logically segment the printable area of ​​the entire consumable. The typesetting equipment designates the largest continuous area on the starting plane, with a single orthogonal viewing angle under the expected installation state, as the main viewing area, serving as the area carrying core information. Continuous areas that are deflected at angles due to crossing physical deformation boundaries and are at non-coplanar viewing angles are designated as secondary viewing areas, serving as the area carrying secondary information. Simultaneously, based on the physical specifications of the consumables, the typesetting equipment marks the reserved assembly hole areas and physical gap areas in the coordinate system, classifying them as avoidance areas with structural obstructions, thus providing a spatial constraint map for subsequent intelligent typesetting.

[0044] In some embodiments, physical deformation boundaries and region divisions can be determined based on original captured images and preset application scenario configuration instructions in a variety of ways:

[0045] Optionally, the typesetting device extracts background environment image features located outside the document boundary from the original captured image, and performs edge gradient detection on the grayscale matrix corresponding to the background environment image features to determine the binary image of the environment edge. Then, it performs Hough line transform on the binary image of the environment edge with preset angle resolution and preset distance resolution to determine multiple sets of parametric lines with accumulator peak values ​​greater than preset truncation thresholds. Subsequently, it determines parallel contour boundary line segments with parallelism errors lower than the allowable threshold and exceeding the document boundary from the multiple sets of parametric lines. Then, it makes orthogonal projection rays from the vertices of the document boundary to the parallel contour boundary line segments. If it is determined that the topological closed region enclosed by the orthogonal projection rays and the parallel contour boundary line segments includes the document boundary, the topological closed region is used as the two-dimensional orthogonal projection region of the target physical carrier. It extracts the light and shadow gradient vectors of the background pixels in the two-dimensional orthogonal projection region and combines them with the contour distortion rate of the topological closed region to form a region structure descriptor. It performs feature space matching between the region structure descriptor and the preset geometric primitive library to determine the geometric solid with the highest confidence interval as the target three-dimensional topological shape, and obtains the preset The application scenario configuration command carries the reference physical size. Based on the first pixel span value of the document boundary in the original captured image, the mapping ratio coefficient between the reference physical size and the first pixel span value is calculated. Based on the second pixel span value of the parallel contour boundary line segment in the original captured image, the product of the mapping ratio coefficient and the second pixel span value is determined as the estimated cross-sectional size. A virtual three-dimensional folding model matching the target three-dimensional topology is determined from the preset consumable folding mapping library. The estimated cross-sectional size and the two-dimensional geometric size of the consumable to be printed are substituted into the virtual three-dimensional folding model to simulate and determine the projection coordinates of the virtual three-dimensional folding model on the two-dimensional unfolded plane. The projection coordinates are used as the physical deformation boundary of the consumable to be printed under the target topology. Finally, the continuous area in the printable area of ​​the consumable to be printed that is in the starting plane and has a single orthogonal viewing angle is divided into the main viewing area using the physical deformation boundary as the dividing line. The continuous area that is deflected at an angle due to crossing the physical deformation boundary is divided into the auxiliary viewing area. The reserved assembly hole area and physical gap area on the consumable to be printed are divided into the avoidance area.

[0046] Optionally, the typesetting device can directly obtain the standard 3D model file of the target physical carrier by parsing the RFID tag scanning data carried in the preset application scenario configuration instructions. Then, it uses the built-in physics engine to perform finite element collision simulation on the material flexibility parameters of the standard 3D model file and the material to be printed, calculates the stress distribution map when the material to be printed is attached to the surface of the standard 3D model, and then extracts the center line of the connected region where the stress value exceeds the yield threshold as the physical deformation boundary. According to the rate of change of the view normal vector in the standard 3D model file, the region where the rate of change of the normal vector is lower than the smoothing threshold is mapped as the main view area, and the curved transition region where the rate of change of the normal vector is higher than the smoothing threshold is mapped as the secondary view area. At the same time, the avoidance area is directly generated by reading the metadata of the forbidden area in the standard 3D model file.

[0047] It is understandable that other methods can be used to achieve 3D shape prediction and region division, and this is not limited here. It should be added that when dividing the main view area and the secondary view area, the typesetting device will also calculate the area weight of each area. If the area of ​​the secondary view area is too small to accommodate any valid object, it will be temporarily marked as a secondary avoidance area to prevent objects from being forcibly stuffed into it during subsequent typesetting, which would cause printing to exceed the boundary.

[0048] S104. Traverse the mapping space coordinate parameters of each object in the structured data list, calculate the spatial adjacency between two objects, and identify multiple objects with spatial adjacency higher than a preset threshold as semantically connected clusters.

[0049] Spatial adjacency refers to the combined closeness of physical distance and logical semantics between any two graphic / text objects in a two-dimensional layout space. Semantic adhesion clusters refer to a set of multiple objects whose spatial adjacency exceeds a preset threshold after calculation and judgment. These clusters represent graphic / text modules that must be moved and scaled as a whole during layout, such as a warning icon and the warning text directly below it. Spatial physical distance refers to the absolute geometric distance between two objects in a coordinate system, typically expressed as the shortest Euclidean distance between bounding boxes or the distance between their centers. A type association matrix is ​​a predefined two-dimensional numerical table where rows and columns correspond to different preset object types. The elements in the matrix define the logical dependence between different preset object types, representing prior knowledge such as a high correlation between "icons" and "text," while a low correlation exists between "barcodes" and "company logos." An object topology graph is a graph data structure with objects as nodes and relationships between objects as edges, used to represent the networked connection state of all objects in the entire layout. The maximum connected subgraph is the largest set of nodes in an object topology graph where every two nodes are connected by a path and no other node can be added to maintain connectivity.

[0050] After completing the regional division of the target consumables, to prevent the erroneous dispersal of logically related graphic elements into different visible areas during subsequent layout, the layout device performs semantic clustering and adhesion cluster generation steps. Specifically, the layout device first traverses the mapped spatial coordinate parameters in the initial layout data (i.e., after mapping each object in the structured data list), using a double loop to pair them up and calculate the spatial physical distance between any two object bounding boxes. To more accurately reflect visual proximity, the layout device calculates not only the center point distance but also the shortest Manhattan distance between the edges of the two bounding rectangles. Next, the layout device extracts the type identifiers of the two paired objects and uses them as an index to query a pre-set type association matrix to obtain the corresponding semantic association coefficient. This coefficient reflects the probability of these two types of objects appearing together in human cognitive habits. Subsequently, the layout device normalizes the calculated spatial physical distance using a Gaussian decay function or an exponential decay function, converting it into a distance score. This distance score is then weighted and multiplied with the queried semantic association coefficient to obtain the comprehensive spatial adjacency between the two objects. After calculating the spatial adjacency of all object pairs, the typesetting device constructs an initial undirected graph using all objects in the structured data list as nodes. The device compares the spatial adjacency of each object pair with a preset threshold. If the spatial adjacency is higher than the threshold, the two objects are considered visually and logically inseparable, and an undirected edge is created between these two corresponding nodes, thus constructing a complete object topology graph. Finally, the typesetting device applies a depth-first search or breadth-first search algorithm from graph theory to traverse the object topology graph and extract the maximum connected subgraphs. Each extracted maximum connected subgraph has nodes (objects) that are highly interconnected. The typesetting device encapsulates each such maximum connected subgraph and identifies it as an independent semantically connected cluster. In subsequent typesetting adjustments, these clusters will be treated as single rigid entities.

[0051] In some embodiments, the process of calculating spatial adjacency and determining semantically connected clusters can be implemented in a variety of ways: Optionally, the typesetting device traverses the mapped spatial coordinate parameters of each object in the structured data list to calculate the spatial physical distance between two objects, determines the corresponding semantic association coefficient from the preset type association matrix based on the type identifiers of the two objects, calculates the spatial adjacency between the two objects by weighting the spatial physical distance using the semantic association coefficient, then constructs an object topology graph by establishing undirected edges between two nodes with spatial adjacency higher than a preset threshold, using all objects in the structured data list as nodes, and finally extracts each maximum connected subgraph in the object topology graph and determines each maximum connected subgraph as a semantically connected cluster. Optionally, the typesetting device utilizes a graph neural network model for dynamic clustering. First, it encodes the coordinates, size, and type identifier of each object into node feature vectors. Then, it constructs an initial sparse adjacency matrix based on spatial coordinates using the K-nearest neighbor algorithm. Next, it inputs the node feature vectors and the sparse adjacency matrix into a graph attention network. Through a multi-head attention mechanism, it dynamically learns the hidden semantic association weights between nodes, outputting an updated dense adjacency matrix as the spatial adjacency degree matrix. Subsequently, it employs a density-based noise-based spatial clustering algorithm to directly perform clustering analysis on this spatial adjacency degree matrix. Core nodes and boundary nodes belonging to the same density-reachable region are grouped into a semantically cohesive cluster, while isolated noise nodes are marked as discrete objects. It is understood that other methods can also be used to achieve semantic clustering of objects; this is not limited here.

[0052] S105. If it is determined that the projected area of ​​the target semantically attached cluster after being mapped by the global scaling factor crosses the boundary or coverage avoidance area between the main view area and the secondary view area, then calculate the local scaling ratio and offset coordinates of the target semantically attached cluster falling into the main view area or the secondary view area.

[0053] Among them, the target semantic cluster refers to a specific semantic cluster whose overall spatial position conflicts with the physical deformation boundary or avoidance area of ​​the consumable in the preliminary layout data, and is used to represent the set of objects that require special typesetting intervention. The projection area refers to the two-dimensional geometric area occupied by the target semantic cluster in the preset two-dimensional target printing coordinate system, and is used to represent the physical coverage of the cluster.

[0054] After identifying all semantically bounded clusters, the typesetting device first traverses all semantically bounded clusters in the initial layout data, extracts the coordinates of all objects within each cluster, calculates the global bounding rectangle that can enclose the cluster, and uses this rectangle as the projection area of ​​the cluster. Next, the typesetting device performs a geometric intersection test between the coordinates of this projection area and the boundary coordinates of the previously defined main view area, secondary view area, and avoidance area. The typesetting device checks whether the projection area crosses the physical deformation boundary (i.e., fold line) between the main view area and the secondary view area, or whether it overlaps with the avoidance area (i.e., holes or gaps). If the projection area of ​​a target semantically bounded cluster exhibits such crossing or overlapping, it indicates that the cluster will be cut by the fold line or destroyed by the hole in actual use, and the typesetting device triggers a local adjustment mechanism. The typesetting device extracts the initial cluster width, initial cluster height, and initial center point coordinates from the global bounding rectangle of the target semantically bounded cluster. Subsequently, the typesetting device assesses the remaining available space in the main and secondary view areas. If no unoccupied area larger than the global bounding rectangle exists, it selects the largest unoccupied area as the target area and determines its effective width, effective height, and target center point coordinates. To ensure the text and graphics content remains undistorted, under the strict constraint of maintaining the initial cluster width to initial cluster height ratio, the typesetting device calculates a first ratio of the effective width to the initial cluster width and a second ratio of the effective height to the initial cluster height. The typesetting device selects the smaller of the first and second ratios and multiplies it by a preset safety margin coefficient (e.g., 0.95, to prevent content from being too close to the boundary), determining the final result as the local scaling ratio. After virtually scaling the global bounding rectangle according to this local scaling ratio to obtain the target bounding rectangle, the typesetting device calculates the required horizontal and vertical translation distances to align the center point of the target bounding rectangle to the target center point coordinates. The two-dimensional vector containing these two translation distances is then determined as the offset coordinates, preparing transformation parameters for subsequent substantive data updates.

[0055] In some embodiments, the process of calculating the local scaling ratio and offset coordinates can be implemented in a variety of ways: Optionally, the typesetting device extracts the initial cluster width, initial cluster height, and initial center point coordinates based on the global bounding rectangle of the target semantic cluster in the preliminary layout data, determines the target area larger than the global bounding rectangle, and determines the effective area width, effective area height, and target center point coordinates of the target area. Under the constraint of keeping the ratio of the initial cluster width to the initial cluster height unchanged, it calculates the first ratio of the effective area width to the initial cluster width and the second ratio of the effective area height to the initial cluster height. Based on the smaller of the first ratio and the second ratio and the preset safety margin coefficient, the local scaling ratio is determined. After scaling the global bounding rectangle according to the local scaling ratio to obtain the target bounding rectangle, the horizontal translation distance and vertical translation distance required to align the center point of the target bounding rectangle to the target center point coordinates are calculated, and the two-dimensional vector containing the horizontal translation distance and the vertical translation distance is determined as the offset coordinates. Optionally, the typesetting device employs a simulated annealing optimization algorithm based on pixel-level masks. First, the device renders the target semantic clusters as high-resolution two-dimensional binary masks, and then renders the available space of the main and secondary view areas as target masks. Next, it initializes a parameter vector containing a scaling factor and a two-dimensional translation vector. The simulated annealing algorithm iteratively searches the parameter space. In each iteration, an affine transformation is performed on the binary mask based on the current parameter vector, and the intersection-union ratio (IUU) of the transformed mask and the target mask, as well as the distance penalty term from the physical deformation boundary, are calculated to construct an energy function. As the system temperature gradually decreases, the algorithm eventually converges to the global optimal solution with the minimum energy function. The typesetting device directly extracts the local scaling ratio and offset coordinates from this optimal solution. It is understood that other methods can also be used to calculate the local adjustment parameters; this is not limited here. It should be noted that when determining the target area, the typesetting device prioritizes evaluating the remaining space in the main view area.

[0056] S106. Update the preliminary layout data according to the local scaling ratio and offset coordinates to obtain the final layout data.

[0057] The cluster placeholder outline area refers to the new rectangular boundary actually occupied by the target semantically attached cluster in the target printing coordinate system after applying local scaling and offset coordinates, representing the absolute sphere of influence of the cluster after adjustment. Discrete objects refer to independent graphic elements that were not classified into any semantically attached clusters during the semantic clustering stage, or belong to other attached clusters that do not require local adjustment, representing relatively free layout blocks. Affected discrete objects refer to discrete objects whose original positions overlap with the updated cluster placeholder outline area due to changes in the position and size of the target semantically attached cluster, representing graphic elements that need to be passively avoided. The joint obstacle boundary is a set of inviolable geometric polygons formed by the updated cluster placeholder outline area and the system-preset avoidance area, representing absolute no-go zones in the layout space. The blank projection point refers to the safe coordinate position outside the joint obstacle boundary that is closest to the original position of the affected discrete object and can completely accommodate the object.

[0058] After calculating the local adjustment parameters of the target semantically bounded cluster, the typesetting device first performs synchronous spatial adjustments on all objects belonging to the target semantically bounded cluster in the preliminary layout data based on the calculated local scaling ratio and offset coordinates. The typesetting device multiplies the coordinates and size of each object within the cluster by the local scaling ratio and adds the offset coordinates, thereby safely migrating the entire cluster to the target area and generating an updated cluster placeholder bounding box area. However, this local migration is highly likely to "hit" other graphic elements originally located within the target area. Therefore, the typesetting device then iterates through all discrete objects in the preliminary layout data that do not belong to the target semantically bounded cluster. The typesetting device extracts the mapped spatial coordinate parameters of these discrete objects and performs a strict Boolean intersection operation with the updated cluster placeholder bounding box area. If the intersection area of ​​the operation result is greater than zero, the typesetting device marks the discrete object as an affected discrete object with overlapping positions. To find new places for these affected discrete objects, the typesetting device merges the updated cluster placeholder bounding box area with the original avoidance area, constructing a complex joint obstacle boundary. Starting with the mapped spatial coordinates of the affected discrete objects, the layout device employs an outward-radiating heuristic search algorithm to find a blank projection point in the available space outside the joint obstacle boundary that is shortest to the starting point and has an area sufficient to accommodate the object. Once found, the layout device generates secondary offset coordinates to move the affected discrete objects to the corresponding blank projection points and uses these secondary offset coordinates to correct the positions of the affected discrete objects in the initial layout data. Finally, the layout device integrates and reorganizes the spatially adjusted data of the target semantic clusters with the position-corrected data of all discrete objects, discarding old coordinate information to generate the final layout data.

[0059] In some embodiments, the process of updating the preliminary layout data to obtain the final layout data can be implemented in multiple ways: Optionally, the typesetting device performs spatial adjustment on the objects contained in the target semantic cluster in the preliminary layout data based on the local scaling ratio and offset coordinates to generate an updated cluster placeholder outline area. It traverses the discrete objects in the preliminary layout data that do not belong to the target semantic cluster and performs Boolean intersection operation between the mapped spatial coordinate parameters of the discrete objects and the updated cluster placeholder outline area to determine the affected discrete objects with overlapping positions. Using the updated cluster placeholder outline area and the avoidance area as the joint obstacle boundary, it determines the blank projection point with the shortest distance to the mapped spatial coordinate parameters of the affected discrete objects. It generates secondary offset coordinates to move the affected discrete objects to the corresponding blank projection points and uses the secondary offset coordinates to correct the position of the affected discrete objects in the preliminary layout data. Finally, it reassembles the spatially adjusted data of the target semantic cluster with the data of all discrete objects after position correction to obtain the final layout data. Optionally, the typesetting device employs a global dynamic obstacle avoidance algorithm based on physical force field simulation. The device sets the updated cluster occupant outline and avoidance zone as high-potential-energy obstacles with a strong repulsive force field, and sets all discrete objects as charged rigid particles. Then, a time-stepping simulation is initiated in a virtual two-dimensional physics engine. Affected discrete objects automatically slide towards empty areas with lower potential energy under the influence of repulsive forces. Simultaneously, weak repulsive forces exist between discrete objects to prevent mutual compression. The simulation stops when the total kinetic energy of the system decays to near zero and all objects are outside the joint obstacle boundary. The typesetting device extracts the final stationary coordinates of each discrete object at this point, calculates the secondary offset coordinates, and reassembles them to generate the final layout data. It is understood that other methods can also be used to update and reassemble the layout data; this is not limited here.

[0060] In this embodiment, by employing physical deformation boundary prediction based on the target's three-dimensional topology and local adaptive typesetting technology for semantically connected clusters, the typesetting device can perceive the spatial folding and surface extension state of the consumables during actual pasting in advance during the two-dimensional typesetting stage. Furthermore, when it detects a set of strongly logically related graphic objects (i.e., semantically connected clusters) crossing the visible area boundary or covering the avoidance area, it breaks the global scaling limit and performs independent local scaling and offset adjustments. This process effectively solves the technical problem in existing technologies where relying solely on two-dimensional geometric collision rules for flow avoidance leads to the unintended allocation of originally closely connected related graphic elements to both sides of the label's fold line or surface extension line. This significantly improves the accuracy and visual integrity of graphic information transmission on special label carriers in complex application scenarios.

[0061] In the above embodiments, the typesetting device can achieve the aforementioned typesetting optimization effect by integrating the technical features of 3D topology prediction and semantic cluster local adjustment. However, in practical applications, when executing the above-mentioned on-demand typesetting method, it often faces situations where the text and image layout in the original captured image is extremely dense and complex. If only macroscopic spatial adjacency determination is relied upon, it is easy to overlook the inherent logical dependencies between different text and image types. This leads to objects that are physically close but semantically unrelated (such as independent barcodes and adjacent decorative borders) being incorrectly clustered into the same semantically connected cluster, resulting in distortion of local scaling ratio calculation and serious waste of typesetting space. The on-demand typesetting method can solve this technical problem by integrating the weighted calculation of semantic association coefficients based on the type association matrix and the extraction of the maximum connected subgraph of the object topology graph, thereby improving the accuracy of semantically connected cluster construction and further enhancing the rationality of the final printed layout and the professionalism of information presentation.

[0062] Please see Figure 2 This is another flowchart illustrating a method for typesetting by shooting and printing in an embodiment of this application.

[0063] S201. Input the original captured image into the preset image recognition model and output a structured data list containing multiple preset object types.

[0064] S202. After determining the global scaling factor of the original captured image in the preset two-dimensional target printing coordinate system based on the document boundary, the global scaling factor is used to map the mapping space coordinate parameters of each object in the structured data list to the two-dimensional target printing coordinate system to generate preliminary layout data.

[0065] S203. Based on the original captured image and the preset application scenario configuration instructions, determine the physical deformation boundary of the consumable to be printed under the target three-dimensional topological shape, and divide the printable area of ​​the consumable to be printed into the main viewing area and the auxiliary viewing area at non-coplanar viewing angles, as well as the avoidance area where there is structural obstruction, based on the physical deformation boundary.

[0066] Steps S201~S203 and Figure 1 The steps S101 to S103 in the illustrated embodiment are similar, and can be referred to the descriptions in steps S101 to S103, which will not be repeated here.

[0067] S204. Traverse the mapping space coordinate parameters of each object in the structured data list and calculate the physical distance between the two objects.

[0068] The structured data list refers to a dynamic array generated by the typesetting equipment during the image recognition stage, containing attribute information for all graphic and text elements. It represents all independent visual entities to be processed in the current typesetting task. An object is the basic data unit in the structured data list, representing a single graphic or text element identified in an image, such as a piece of text, a barcode, or a company logo. Mapped spatial coordinate parameters refer to the geometric position and size boundary data of each object in a two-dimensional plane coordinate system, representing the coordinates of the top-left vertex, absolute width, and absolute height of the object's bounding rectangle. Spatial physical distance refers to the absolute length measure between the bounding rectangles or geometric center points of any two objects in two-dimensional geometric space, representing the visual proximity of these two graphic and text elements.

[0069] After completing the 3D topology prediction and spatial dimensionality reduction partitioning of the target consumables, in order to maintain the original visual layout logic of the graphics and text in subsequent typesetting, it is necessary to quantitatively evaluate the geometric proximity relationships of all discrete graphics and text elements. In this scenario, the typesetting device performs the step of traversing coordinate parameters and calculating spatial physical distances. Specifically, the typesetting device first reads a structured data list from memory and initializes a two-dimensional distance matrix or hash map table to store the distance calculation results. The typesetting device uses a double-nested loop logic structure to fully traverse all objects in the list. In the outer loop, a reference object is locked, and in the inner loop, the remaining comparison objects are extracted sequentially. For each selected pair of reference and comparison objects, the typesetting device parses their mapped spatial coordinate parameters and extracts the coordinate values ​​of the left, right, top, and bottom boundaries of their respective bounding rectangles. To comprehensively and accurately reflect the true visual distance between the two objects, the typesetting device not only calculates the Euclidean distance between the geometric center points of the two objects but also calculates the shortest Manhattan distance between the edges of their bounding rectangles. The typesetting device determines whether two rectangles have projected overlap in the horizontal or vertical directions by comparing their relative positions, such as the right boundary of the reference object and the left boundary of the comparison object, and the bottom boundary of the reference object and the top boundary of the comparison object. If projected overlap exists, the typesetting device sets the edge distance in the corresponding direction to zero; if no overlap exists, the typesetting device calculates the absolute difference between the nearest edges. Subsequently, the typesetting device linearly merges the Euclidean distance of the center point and the shortest edge distance according to a preset weight ratio to obtain a comprehensive scalar value. This value is determined as the spatial physical distance between the two objects, and the distance value, along with the unique identifiers of the two objects, is stored in a pre-initialized two-dimensional distance matrix until all object pairs have been traversed and calculated.

[0070] In some embodiments, the process of traversing coordinate parameters and calculating spatial physical distances can be implemented in several ways: Optionally, the typesetting device employs accelerated calculation logic based on a spatial index tree. The typesetting device first reads the mapped spatial coordinate parameters of all objects, uses these parameters to construct a quadtree or R-tree spatial index structure in memory, then traverses each object in the structured data list as a query node, and uses the spatial index tree to quickly retrieve all neighboring objects within the preset search radius of the query node. Subsequently, it extracts the boundary coordinates only for these retrieved neighboring objects, calculates the shortest edge distance between them and the query node as the spatial physical distance. This approach effectively avoids the massive computational overhead of global pairwise traversal. Optionally, the typesetting device employs a grid-based distance estimation logic. The device divides the entire two-dimensional target printing coordinate system into a uniform grid of fixed size. Based on the mapped spatial coordinate parameters of each object, it maps it to the corresponding grid cell. Then, it traverses the non-empty grid containing the objects. For any two objects, the typesetting device quickly estimates the grid distance by calculating the difference between the row and column indices of their respective grid cells. Subsequently, for object pairs with a grid distance less than a preset threshold, it extracts their center point coordinates to calculate a high-precision Euclidean distance as the final physical distance. It is understood that other methods can also be used to calculate the physical distance between objects; this is not limited here.

[0071] S205. Based on the type identifiers of the two objects, determine the corresponding semantic association coefficients from the preset type association matrix.

[0072] In this context, "type identifier" refers to the classification attribute label assigned to each text and image object during the image recognition stage. This label indicates the logical semantic type of information carrier the object belongs to, such as "main title," "body description," "warning icon," or "QR code." The pre-set type association matrix is ​​a two-dimensional numerical lookup table pre-loaded in the typesetting device's memory. Its row and column indices correspond to all preset object types supported by the system. Each intersection element in the matrix defines the degree of dependence between two specific object types in conventional typesetting logic, representing prior typesetting rule knowledge. The semantic association coefficient is a specific floating-point value extracted from the type association matrix. It indicates whether the two specific objects currently being evaluated have a strong necessity for combination in terms of content expression; a larger value indicates a stronger logical association.

[0073] After calculating the absolute physical distance of all object pairs, the typesetting device performs semantic association coefficient determination based on type identifiers. Specifically, during the traversal of object pairs, for any two objects currently being processed, the type identifier field bound to them in the structured data list is read. The typesetting device uses the type identifier of the first object as the row search key and the type identifier of the second object as the column search key. Subsequently, the typesetting device accesses a pre-loaded type association matrix in memory, using these two search keys to locate the two-dimensional array. This type association matrix is ​​pre-constructed based on statistical analysis of a large amount of industrial label and instruction manual typesetting data. For example, the matrix defines a very high intersection value between "warning icon" and "warning text" because they are semantically inseparable; while the intersection value between "company logo" and "production batch barcode" is very low because they usually express information independently. After locating the corresponding row and column intersection point in the matrix, the typesetting device reads the floating-point value stored at that location. To ensure data processing symmetry, the typesetting device also verifies whether the matrix is ​​symmetric. If it is asymmetric (i.e., the dependency of A on B is not equal to the dependency of B on A), the typesetting device extracts the values ​​of swapped rows and columns and takes the average or maximum of the two values. Finally, the typesetting device determines the extracted and confirmed values ​​as the semantic association coefficient between the two objects. This coefficient will serve as the key logical multiplier for subsequent adjustments to the physical distance weights and is temporarily stored in the typesetting device's cache register, awaiting participation in the next step of fusion calculation.

[0074] In some embodiments, the process of determining the semantic association coefficient can be implemented in several ways: Optionally, the typesetting device adopts a fast table lookup logic based on static hash mapping. The typesetting device first concatenates the type identifiers of the two objects according to alphabetical order or preset encoding rules to generate a unique combined string key value. Then, it uses the combined string key value to perform hash addressing in a static hash table pre-loaded into memory. If the key value exists in the hash table, the corresponding value is directly extracted as the semantic association coefficient. If the key value does not exist in the hash table (i.e., an undefined type combination is encountered), the typesetting device returns a preset default base coefficient as the semantic association coefficient. Optionally, the typesetting device adopts an adaptive logic based on a dynamic loading matrix of application scenarios. The typesetting device first parses the preset application scenario configuration instructions to identify the industry field to which the current printing task belongs (such as power inspection, medical devices, or warehousing and logistics). Then, according to the identified industry field, it dynamically downloads and loads a specific type association matrix that highly matches the field from local storage or cloud database. Subsequently, it extracts the type identifiers of the two objects, performs row and column cross-indexing in the specific type association matrix, and extracts the value that conforms to the current industry typesetting specification as the semantic association coefficient. It is understandable that other methods can be used to determine the semantic association coefficient, which are not limited here.

[0075] S206. The spatial adjacency between two objects is obtained by weighting the spatial physical distance using semantic association coefficients.

[0076] The typesetting device first synchronously reads the spatial physical distance and semantic association coefficient for the same pair of objects from memory. Since spatial physical distance is a quantity with absolute physical units (such as pixels or millimeters), where smaller values ​​indicate closer proximity, while the semantic association coefficient is a dimensionless quantity where larger values ​​indicate greater relevance, the typesetting device cannot directly multiply them. Therefore, the typesetting device first normalizes and inverts the spatial physical distance. It substitutes the spatial physical distance into a preset monotonically decreasing function (such as a Gaussian decay function or an exponential decay function) to calculate a spatial proximity score between zero and one. When the physical distance is close to zero, the score approaches one; when the physical distance approaches infinity, the score approaches zero. Next, the typesetting device extracts the semantic association coefficient and ensures that its value range is also normalized to between zero and one. Subsequently, the typesetting device executes the core weighted calculation logic. It employs a non-linear multiplication fusion strategy to multiply the calculated spatial proximity score by the semantic association coefficient. This multiplication logic ensures that a high score is only achieved when two objects are physically close enough and logically highly related; if either is extremely low (e.g., close in distance but mutually exclusive in type, or related in type but extremely far apart), the product result will be lowered. Finally, the typesetting device determines this product result as the spatial adjacency between the two objects and updates this value in the previously constructed two-dimensional matrix, replacing the original single physical distance value.

[0077] In some embodiments, the process of weighted calculation and obtaining spatial adjacency can be implemented in multiple ways: Optionally, the typesetting device adopts a calculation logic based on linear weighting and threshold truncation. The typesetting device first uses the maximum-minimum value normalization method to convert the spatial physical distance into a distance penalty term between zero and one. Then, it assigns preset weight ratio parameters to the distance penalty term and the semantic association coefficient respectively. Subsequently, it calculates the product of the semantic association coefficient and the corresponding weight, subtracts the product of the distance penalty term and the corresponding weight, and obtains an initial fusion value. Finally, it uses an activation function to truncate the part of the initial fusion value that is less than zero to zero, and retains the part that is greater than zero and determines it as the final spatial adjacency. The typesetting equipment employs nonlinear fusion logic based on a fuzzy logic inference system. First, the spatial physical distance is input into a preset distance membership function, which is then fuzzified into membership vectors of fuzzy sets such as "very close," "relatively close," and "relatively far." Simultaneously, the semantic association coefficient is input into an association membership function, which is also fuzzified into membership vectors of fuzzy sets such as "strong association" and "weak association." Next, a preset fuzzy inference rule base (e.g., "if the distance is very close and there is a strong association, then the adjacency degree is very high") is used to perform fuzzy intersection and union operations on these two vectors. Finally, the inference result is defuzzified using the centroid method, outputting a continuous value as the spatial adjacency degree. It is understood that other methods can also be used to achieve weighted fusion of multi-dimensional indicators; this is not limited here. It should be noted that during the weighted calculation, the typesetting device also introduces a directional penalty factor. The typesetting device calculates the angle of the line connecting the center points of two objects. If this angle deviates from the standard horizontal or vertical alignment direction (for example, presenting a 45-degree diagonal arrangement), the typesetting device will multiply the final spatial adjacency by a penalty coefficient less than one, because in regular tag typesetting, diagonally arranged objects usually do not belong to the same semantic group.

[0078] S207. Using all objects in the structured data list as nodes, establish undirected edges between two nodes whose spatial adjacency is higher than a preset threshold to construct an object topology graph.

[0079] In this context, a node refers to the basic vertex element in a graph data structure, representing each independent text / graphic object in a structured data list. A preset threshold is a numerical limit set in the typesetting device's memory for binary classification, representing the minimum standard for determining whether two objects have a sufficiently strong comprehensive correlation to be bound together. An undirected edge is a non-directional mathematical connection between two nodes, indicating a bidirectional, inseparable, and strongly bound relationship between the objects corresponding to these nodes in the typesetting logic. An object topology graph is a non-linear data structure composed of all nodes and the undirected edges connecting them, representing the global networked connection state of all text / graphic elements in the entire two-dimensional typesetting layout.

[0080] After completing the spatial adjacency metric calculation for all object pairs, the typesetting device executes the step of constructing an object topology graph. Specifically, the typesetting device first initializes an empty graph data structure in memory, typically in the form of an adjacency list or adjacency matrix. The typesetting device traverses the structured data list, instantiating each object in the list as an independent node, and attaching the object's unique identifier, mapping spatial coordinate parameters, and type identifier as node attributes to that node, thereby injecting all objects into the graph data structure. Next, the typesetting device reads a pre-set threshold, which determines the sensitivity of the clustering algorithm. The typesetting device traverses the previously generated two-dimensional matrix containing the spatial adjacency degrees of all object pairs. For each spatial adjacency value in the matrix, the typesetting device rigorously compares it with the preset threshold. If a spatial adjacency degree is lower than or equal to the preset threshold, the typesetting device determines that the association between the two objects is insufficient to form a rigid binding, and therefore ignores the relationship between the two nodes in the graph; if a spatial adjacency degree is significantly higher than the preset threshold, the typesetting device determines that the two objects are visually and semantically highly connected and inseparable. In this scenario, the typesetting device locates the nodes corresponding to the two objects within the graph data structure, instantiates them, and inserts an undirected edge between them. The device also assigns the specific value of the spatial adjacency as a weight attribute to this undirected edge to preserve information about the strength of the association. As the traversal progresses, the typesetting device continuously adds undirected edges between nodes that meet the criteria, eventually constructing a complete object topology graph in memory. In this graph, closely related objects are connected by dense edges, forming a network, while isolated objects are represented as isolated nodes without any connecting edges.

[0081] In some embodiments, the process of constructing an object topology graph can be implemented in several ways: Optionally, the typesetting device adopts a static graph construction logic based on an adjacency matrix. First, based on the total number N of objects in the structured data list, the typesetting device allocates an N x N two-dimensional Boolean adjacency matrix in memory and initializes all values ​​to false. Then, it iterates through the numerical matrix containing spatial adjacency degrees. When the extracted spatial adjacency degree is greater than a preset threshold, the typesetting device modifies the values ​​of the corresponding row and column intersections and the symmetrical row and column intersections in the Boolean adjacency matrix to true, thereby implicitly establishing an undirected edge between the two nodes. Finally, the... Boolean adjacency matrices serve as the underlying data representation of the object topology graph. Optionally, the typesetting device employs a graph construction logic based on dynamic adjacency lists. The device initializes an empty linked list as its neighbor list for each object in the structured data list. Then, it traverses the spatial adjacency data. When two objects are found to have a spatial adjacency higher than a preset threshold, the device encapsulates the node pointers and adjacency weights of the other object into edge objects and dynamically inserts them into each other's neighbor lists. This approach significantly saves memory space and improves the efficiency of subsequent graph traversal when handling typesetting tasks with a large number of objects and sparse connections. It is understood that other methods can also be used to construct the graph data structure; this is not limited here. It should be noted that when establishing undirected edges, the typesetting device also performs intersection detection. If a newly established undirected edge intersects with other existing undirected edges in the two-dimensional geometric projection, the device re-evaluates the spatial adjacency of these two pairs of objects and may forcibly disconnect the edge with the lower weight to ensure the logical rationality of the generated object topology graph in planar typesetting and avoid visual logical confusion.

[0082] In some embodiments, the typesetting device introduces data processing logic based on local density adaptive thresholds and K-nearest neighbor verification. Instead of using a single globally preset threshold, the typesetting device calculates the average and standard deviation of the adjacency of all objects within a certain physical radius around each node, dynamically generating a local adaptive threshold for that node. Only when the spatial adjacency between two objects is simultaneously higher than their respective local adaptive thresholds are they initially qualified for an edge connection. Furthermore, the typesetting device extracts the set of the K objects with the highest spatial adjacency of each node (i.e., the K-nearest neighbor set). When deciding whether to establish an undirected edge between nodes A and B, the typesetting device not only requires their adjacency to meet the threshold but also strictly requires that node A must exist in node B's K-nearest neighbor set, and node B must also exist in node A's K-nearest neighbor set. Through this bidirectional K-nearest neighbor verification mechanism, the typesetting device effectively cuts off unidirectional weak association transmission, completely suppresses chain effects, and ensures that the connectivity relationships in the constructed object topology graph have extremely high local cohesion.

[0083] S208. Extract each maximum connected subgraph from the object topology graph and identify each maximum connected subgraph as a semantically connected cluster.

[0084] In graph theory, a maximum connected subgraph refers to a local network in an object topology graph consisting of a set of nodes and their undirected edges. In this local network, there exists at least one path between any two nodes, and this fully connected state cannot be maintained by introducing any other node. Semantic clusters, on the other hand, are rigid graph modules formed after extraction and encapsulation using graph theory algorithms. They represent a set of objects that must undergo synchronous geometric transformations as an indivisible whole during subsequent local scaling and offset adjustments.

[0085] In scenarios where the networked node connections need to be transformed into concrete, rigid modules usable for typesetting operations after the object topology graph has been constructed, the typesetting device performs the steps of extracting the maximum connected subgraph and determining semantically connected clusters. Specifically, the typesetting device first initializes a Boolean flag array in memory to record the node visit status, initializing the state of all nodes to "unvisited," and simultaneously initializing an empty list to store the final generated semantically connected clusters. Next, the typesetting device traverses all nodes in the object topology graph in node index order. When a starting node with a state of "unvisited" is encountered, the typesetting device initiates a graph traversal algorithm (e.g., depth-first search or breadth-first search) starting from that node. During the traversal, the typesetting device continuously visits adjacent nodes along undirected edges, updating the state of each visited node to "visited," and continuously adding these nodes to a temporary subgraph set. The traversal process continues to expand outward until all adjacent nodes on the current connected path have been visited, and no new connected nodes can be found. At this point, all the nodes collected in the temporary subgraph set constitute a maximally connected subgraph that strictly satisfies the requirements of maximization and connectivity. The typesetting device then objectifies and encapsulates this maximally connected subgraph. It extracts the mapping space coordinate parameters of all text and image objects contained in the subgraph, calculates the global bounding rectangle that completely encloses these objects, and uses the size and coordinates of this global bounding rectangle as the overall physical boundary of the cluster. The typesetting device defines the encapsulated data structure containing multiple objects and their overall boundary information as an independent semantically connected cluster and pushes it into a pre-initialized list. The typesetting device continues to traverse the remaining "unvisited" nodes in the marker array, repeating the above search and encapsulation process until every node in the graph has been visited and categorized. Finally, the typesetting device successfully decomposes the entire complex object topology graph into multiple independent semantically connected clusters, providing clear operational entities for subsequent conflict detection and local typesetting adjustments.

[0086] In some embodiments, the process of extracting the maximum connected subgraph and determining semantically connected clusters can be implemented in several ways: Optionally, the typesetting device adopts a recursive extraction logic based on depth-first search. The typesetting device defines a recursive function that takes the current node and the current subgraph set as parameters. Inside the function, the typesetting device marks the current node as visited and adds it to the subgraph set. Then, it traverses all adjacent nodes of the current node. If an adjacent node has not been visited, the function is recursively called with that adjacent node as a parameter. Through this deep recursive backtracking, the typesetting device can completely mine all nodes of a maximum connected subgraph, and then calculate the coordinate extrema of these nodes to generate the outer moments. The graph is encapsulated and identified as a semantically connected cluster. Optionally, the typesetting device employs dynamic connected component extraction logic based on a disjoint-set data structure. First, the typesetting device initializes an independent set for each node in the graph. Then, it traverses all undirected edges in the object's topology graph. For each pair of nodes connected by an edge, the typesetting device performs a disjoint-set merge operation, merging their sets into a larger set. Simultaneously, path compression technology is used to optimize search efficiency. After all edges have been traversed, each remaining independent set in the disjoint-set corresponds to a maximum connected subgraph. The typesetting device directly extracts the node data from these sets, calculates the overall boundary, and identifies it as a semantically connected cluster. It is understood that other methods can also be used to extract and encapsulate connected subgraphs; this is not limited here. It should be added that after determining the semantically connected clusters, the typesetting device will also perform a size verification on the generated clusters. If a certain extracted maximum connected subgraph contains only a single node (i.e., the node has no connecting edges in the graph), the typesetting device will not encapsulate it as a semantically connected cluster, but will mark it as a discrete object, so that it follows the conventional streaming avoidance rules in subsequent typesetting, thereby saving the computational resources for local adjustments.

[0087] S209. If it is determined that the projected area of ​​the target semantically attached cluster after being mapped by the global scaling factor crosses the boundary or coverage avoidance area between the main view area and the secondary view area, then calculate the local scaling ratio and offset coordinates of the target semantically attached cluster falling into the main view area or the secondary view area.

[0088] S210. Update the preliminary layout data according to the local scaling ratio and offset coordinates to obtain the final layout data.

[0089] Steps S209~S210 and Figure 1 Steps S105-S106 in the illustrated embodiment are similar and can be found in the descriptions of steps S105-S106, which will not be repeated here.

[0090] In this embodiment, the fusion technology based on the semantic association coefficient of the type association matrix and the weighted calculation of spatial physical distance is adopted to quantify the comprehensive tightness between objects. This effectively solves the technical problem in the prior art of incorrectly binding logically unrelated graphic elements (such as independent barcodes and adjacent decorative borders) based solely on their physical proximity, resulting in wasted layout space and local scaling distortion. This improves visual integrity and the professionalism of information transmission.

[0091] The exemplary typesetting device 300 provided in the embodiments of this application is described below. Figure 3 This is an exemplary hardware structure diagram of the typesetting device 300 provided in the embodiments of this application.

[0092] In some embodiments, the typesetting device 300 is a computer device or includes a computer device. The computer device includes a processor, memory, and a network interface connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores data. The network interface of the computer device is used to communicate with other external terminals or servers via a network connection. In some embodiments, the network interface can be a wired network interface; in some embodiments, the network interface can also be a wireless network interface. When the computer program is executed by the processor, it implements the methods in the embodiments of this application.

[0093] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0094] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0095] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as meaning "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is interpreted as meaning "if determining...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".

[0096] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.

[0097] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A method for typesetting images on the fly, characterized in that, Applied to typesetting equipment, the method includes: The original captured image is input into a preset image recognition model, and the output is a structured data list containing multiple preset object types. The objects in the structured data list include corresponding source image data, corresponding type identifiers, and mapping space coordinate parameters in the original captured image. After determining the global scaling factor of the original captured image in the preset two-dimensional target printing coordinate system based on the document boundary, the mapping space coordinate parameters of each object in the structured data list are mapped to the preset two-dimensional target printing coordinate system using the global scaling factor to generate preliminary layout data. The document boundary is determined based on the object distribution characteristics in the structured data list. Based on the original captured image and the preset application scenario configuration instructions, the physical deformation boundary of the consumable to be printed under the target three-dimensional topological shape is determined, and based on the physical deformation boundary, the printable area of ​​the consumable to be printed is divided into the main viewing area and the auxiliary viewing area at non-coplanar viewing angles, as well as the avoidance area where there is structural obstruction. Traverse the mapping space coordinate parameters of each object in the structured data list, calculate the spatial adjacency between two objects, and identify multiple objects with spatial adjacency higher than a preset threshold as semantically connected clusters; If it is determined that the projection area of ​​a target semantically connected cluster after being mapped by the global scaling factor crosses the boundary between the main view area and the secondary view area or covers the avoidance area, then the local scaling ratio and offset coordinates of the target semantically connected cluster falling into the main view area or the secondary view area are calculated. The final layout data is obtained by updating the preliminary layout data according to the local scaling ratio and the offset coordinates.

2. The method according to claim 1, characterized in that, The step of determining the physical deformation boundary of the filament to be printed under the target three-dimensional topological shape based on the original captured image and the preset application scenario configuration instruction, and dividing the printable area of ​​the filament to be printed into a main viewing area and a secondary viewing area at non-coplanar viewing angles, as well as an avoidance area with structural obstructions, based on the physical deformation boundary, specifically includes: Based on the background environment image features located outside the document boundary in the original captured image, the target three-dimensional topological shape and estimated cross-sectional size of the target physical carrier corresponding to the document are determined; Determine a virtual 3D folding model that matches the target 3D topology from a pre-set consumable folding mapping library; The estimated cross-sectional dimensions and the two-dimensional geometric dimensions of the consumable to be printed are substituted into the virtual three-dimensional folding model for simulation. The projection coordinates of the virtual three-dimensional folding model on the two-dimensional unfolding plane are determined, and the projection coordinates are used as the physical deformation boundary of the consumable to be printed under the target three-dimensional topological shape. Using the physical deformation boundary as a dividing line, the continuous area in the printable area of ​​the consumable to be printed, which is located on the starting plane and has a single orthogonal viewing angle, is divided into the main viewing area. The continuous area that is deflected at an angle due to crossing the physical deformation boundary is divided into the auxiliary viewing area. The pre-reserved assembly hole area and physical gap area on the consumable to be printed are divided into the avoidance area.

3. The method according to claim 2, characterized in that, The step of determining the target three-dimensional topological shape and estimating the cross-sectional size of the target physical carrier corresponding to the document based on the background environment image features located outside the document boundary in the original captured image specifically includes: Extract background environment image features located outside the document boundary from the original captured image, and perform edge gradient detection on the grayscale matrix corresponding to the background environment image features to determine the binary image of the environment edge; Perform Hough line transform on the binary map of the environmental edge with preset angle resolution and preset distance resolution to determine multiple sets of parameterized lines whose accumulator peak value is greater than a preset truncation threshold. From the multiple sets of parametric lines, determine the parallel contour boundary line segments whose parallelism error is below the allowable threshold and exceeds the document boundary; Based on the contour deformation distortion rate of the closed region enclosed by the parallel contour boundary line segments, the closed region is mapped to a preset geometric primitive library for matching to determine the target three-dimensional topological shape; After obtaining the reference physical size carried in the preset application scenario configuration instruction, the mapping ratio coefficient between the reference physical size and the first pixel span value is calculated based on the first pixel span value of the document boundary in the original captured image. The reference physical size is the one-dimensional absolute length of the document known in the real world. Based on the second pixel span value of the parallel contour boundary line segment in the original captured image, the product of the mapping ratio coefficient and the second pixel span value is determined as the estimated cross-sectional size.

4. The method according to claim 3, characterized in that, The step of mapping the closed region to a preset geometric primitive library for matching based on the contour deformation distortion rate of the closed region enclosed by the parallel contour boundary line segments to determine the target three-dimensional topological shape specifically includes: After orthogonally projecting rays from the vertices of the document boundary to the parallel contour boundary line segments, if it is determined that the topological closed region enclosed by the orthogonal projection rays and the parallel contour boundary line segments contains the document boundary, the topological closed region is taken as the two-dimensional orthogonal projection region of the target physical carrier. Extract the light and shadow gradient vectors of the background pixels within the two-dimensional orthographic projection region, and combine them with the contour distortion rate of the topological closure region to form a region structure descriptor; The region structure descriptor is matched with the preset geometric primitive library to determine the geometric solid with the highest confidence interval as the target three-dimensional topological shape.

5. The method according to claim 1, characterized in that, The step of calculating the local scaling ratio and offset coordinates of the target semantic adhesion cluster falling into the main view area or the secondary view area specifically includes: Based on the global bounding rectangle of the target semantic cluster in the preliminary layout data, the initial cluster width, initial cluster height, and initial center point coordinates are extracted from the global bounding rectangle. Determine a target region larger than the global bounding rectangle, and determine the effective area width, effective area height, and target center point coordinates of the target region. The target region is the unoccupied main view area or the secondary view area. Under the constraint of keeping the ratio of the initial cluster width to the initial cluster height constant, calculate a first ratio of the effective area width to the initial cluster width, and a second ratio of the effective area height to the initial cluster height; The local scaling ratio is determined based on the smaller of the first ratio and the second ratio and a preset safety margin coefficient; After scaling the global bounding rectangle according to the local scaling ratio to obtain the target bounding rectangle, calculate the required horizontal and vertical translation distances when aligning the center point of the target bounding rectangle to the coordinates of the target center point, and determine the two-dimensional vector containing the horizontal and vertical translation distances as the offset coordinates.

6. The method according to claim 1, characterized in that, The step of traversing the mapped spatial coordinate parameters of each object in the structured data list, calculating the spatial adjacency between two objects, and identifying multiple objects with spatial adjacency higher than a preset threshold as a semantically cohesive cluster specifically includes: Iterate through the mapped spatial coordinate parameters of each object in the structured data list to calculate the spatial physical distance between two objects; Based on the type identifiers of the two objects, the corresponding semantic association coefficients are determined from the preset type association matrix. The elements in the type association matrix define the degree of logical dependence between different preset object types. The spatial adjacency degree between the two objects is obtained by weighting the spatial physical distance using the semantic association coefficient. Using all objects in the structured data list as nodes, an undirected edge is established between two nodes whose spatial adjacency is higher than a preset threshold to construct an object topology graph; Extract each of the maximum connected subgraphs in the object topology graph, and determine each of the maximum connected subgraphs as the semantic adhesion cluster.

7. The method according to claim 1, characterized in that, The step of updating the preliminary layout data according to the local scaling ratio and the offset coordinates to obtain the final layout data specifically includes: Based on the local scaling ratio and the offset coordinates, spatial adjustment is performed on the objects contained in the target semantic cluster in the preliminary layout data to generate an updated cluster placeholder outline area. Traverse the discrete objects in the preliminary layout data that do not belong to the target semantic adhesion cluster, perform a Boolean intersection operation between the mapping space coordinate parameters of the discrete objects and the updated cluster occupant outline area, and determine the affected discrete objects with overlapping positions. Using the updated cluster occupancy frame region and the avoidance area as the joint obstacle boundary, determine the blank projection point with the shortest distance to the mapped space coordinate parameters of the affected discrete object; Generate secondary offset coordinates to move the affected discrete object to the corresponding blank projection point, and use the secondary offset coordinates to correct the position of the affected discrete object in the preliminary layout data; The data after adjusting the target semantic cluster space and the data after correcting the positions of all discrete object data are recombined to obtain the final layout data.

8. A typesetting device, characterized in that, The typesetting device includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the typesetting device to perform the method as described in any one of claims 1-7.

9. A computer program product containing instructions, characterized in that, When the computer program product is run on a typesetting device, the typesetting device performs the method as described in any one of claims 1-7.

10. A computer-readable storage medium comprising instructions, characterized in that, When the instruction is executed on the typesetting device, the typesetting device performs the method as described in any one of claims 1-7.