Intelligent image-text integrated processing system and method

The intelligent integrated graphic and text processing system solves the problems of cumbersome operation and information loss in traditional graphic and text processing software. It enables seamless processing of vector and raster graphics on a single platform, supports intelligent analysis and automatic generation of professional documents, and improves the efficiency of engineering graphic and text processing.

CN120633589BActive Publication Date: 2026-01-02BEIJING LONGRUAN TECHNOLOGIES INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510722371.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2026-01-02
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

Traditional image processing software is cumbersome to operate when processing vector or raster graphics, requiring multiple software programs to work together. Furthermore, it is easy to lose the spatiotemporal relationships and attribute information of graphics when sharing data across terminals, making it impossible to achieve fast and intelligent analysis and image-to-text conversion. This is particularly inefficient in the engineering field.

Method used

This invention provides an intelligent integrated graphic and text processing system, including a vector graphics processing platform, a graphic and text large model module, a graphic and text layout module, and a graphic and text document generation module. It realizes mixed layout of graphic editing, text processing, and multimedia data through a unified interactive interface, supports manual customization and automated processing, and performs intelligent analysis and generation in conjunction with the graphic and text large model.

Benefits of technology

It enables seamless processing of vector and raster graphics on a single platform, avoiding information loss, supporting intelligent analysis and automatic generation of professional documents, and improving the efficiency and intelligence of engineering graphic processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633589B_ABST
    Figure CN120633589B_ABST
Patent Text Reader

Abstract

The application discloses an intelligent graphic-text integrated processing system and method, and relates to the field of engineering.The system provides a vector graphic processing platform, supports processing of text, vector graphics, raster graphics, animation, video and other multimedia data, and can simultaneously complete processing and mixed layout of the multimedia data without calling external programs.Through construction of a graphic-text large model, intelligent conversion and automatic layout capabilities of text and multimedia data are provided, according to semantic association of the graphic-text, vector / raster graphics automatically generates graphic-text multimedia text description and layout, or through input of natural language, vector / raster graphics is automatically generated, and an intelligent graphic-text special subject document is generated.The application constructs a graphic-text processing and generation large model, and considers interactive, intelligent and automatic graphic-text generation, processing and layout, and is particularly suitable for engineering fields which need a large amount of graphic-text data and have an association relationship between the graphic-text, and replaces traditional text type office software.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of engineering drawings, and in particular to an intelligent drawing-text integrated processing system and method. BACKGROUND

[0002] In daily learning, work and life, drawing-text arrangement software has always played a crucial role and has provided great convenience for people. Traditional drawing-text processing software includes Microsoft's Office series, China's Kingsoft WPS series, etc.; computer-aided design (CAD) software includes AutoCAD, Zhongwang CAD, etc. Geographic information system (GIS) software includes Esri's ArcGIS, SuperMap software's SuperMap, etc. Internet drawing-text processing software includes Baidu and Tencent's intelligent panel, etc.

[0003] The base of traditional drawing-text processing software is a text processing platform and system. When processing vector or raster graphics, it is usually necessary to call external programs (such as Microsoft OLE technology) or use the screenshot method in the text processing system (such as WPS) to complete the processing. This method is cumbersome to operate, requires the operation of multiple software, and after the data is shared across terminals and systems, it may not be able to display or process external graphics normally due to the lack of external calling programs, which has become one of the obstacles affecting the application of drawing-text systems in the Internet era. In addition, the screenshot of vector graphics will lose the description of the spatial relationship and related attribute information of the graphics, which is not conducive to intelligent analysis and drawing-text conversion.

[0004] Current artificial intelligence technology is booming, and traditional office and drawing software can also provide some general AI support (such as style configuration, unified layout, text-to-image or video, etc.). However, due to the lack of key attributes such as spatial relationship and professional drawing-text spatial-temporal large models, especially the lack of intelligent understanding and drawing-text conversion functions for vector graphics, it is not possible to quickly generate professional spatial-temporal intelligent analysis or engineering documents related to graphics in the engineering field, which is mainly completed through interactive manual operation. Professional personnel are inefficient when preparing reports in the field of engineering with strong professional characteristics (such as GIS / CAD engineering papers or reports).

[0005] While the intelligent panel of general Internet companies provides simple vector editing such as graffiti, formulas, flowcharts, mind maps, etc., it lacks drawing-text mixed arrangement, and has no intelligent analysis and processing function for vector engineering graphics. SUMMARY

[0006] In view of the above problems, the present application proposes an intelligent image-text integrated processing system and method to overcome the shortcomings of the prior art.

[0007] The embodiment of the present application provides an intelligent image-text integrated processing system, comprising: a vector graphics processing platform, an image-text large model module, an image-text layout module, and an image-text document generation module.

[0008] The vector graphics processing platform is internally provided with a multimedia data processing engine including text, vector graphics, raster graphics, animation, and video, and is used to directly complete graphic editing, text processing, and mixed layout of multimedia data through a unified graphic interactive interface.

[0009] The image-text layout module is used to provide manual customized image-text processing and layout through graphical processing and interaction based on the vector graphics processing platform, and to provide automatic image-text processing and layout based on the image-text large model module, automatically generate vector graphics or raster graphics and layout corresponding to the content based on the text content according to templates or intelligent analysis, or automatically generate text content and layout corresponding to the vector graphics or raster graphics based on the vector graphics or raster graphics, wherein the templates include templates provided by the image-text large model module or templates provided by a third party.

[0010] The image-text document generation module is used to automatically generate professional image-text document reports based on the image-text layout module, combine given themes or the templates, analyze key attribute data and GIS / CAD spatial relationship data through semantic analysis and correlation analysis, and combine the image-text large model module and the space-time reasoning function.

[0011] Optionally, the vector graphics processing platform comprises an image-text rendering module, an image-text interactive editing module, and an image-text linkage module.

[0012] Through the image-text rendering module, the image-text interactive editing module, and the image-text linkage module, the vector graphics processing platform has the editing capability of vector and raster graphics, and simultaneously provides the arrangement and processing capability of multimedia data on the graphic processing base or platform.

[0013] Through the image-text rendering module, the image-text interactive editing module, and the image-text linkage module, the intelligent image-text integrated processing system is provided with a unified interactive interface, and the editing of vector and raster graphics, the processing of text, and the mixed layout can be completed without calling external programs, seamlessly switching, mutually assisting, eliminating the system-level data exchange link, and avoiding the loss of key information including attributes and topological relationships.

[0014] The mixed layout refers to unified operation on vector graphics, raster graphics and the multimedia data, supporting a tiled sequential layout of an office software through a global offset attribute of a graphic primitive, or supporting a stacked placeholder layout through a fixed position attribute of the graphic primitive.

[0015] The unified interactive interface without calling an external program refers to rendering, editing and layout of the multimedia data, and linkage updating, which are respectively completed by the graphic-text rendering module, the graphic-text interactive editing module and the graphic-text linkage module without the aid of an external program.

[0016] Optionally, the graphic-text rendering module is configured to load the vector and raster graphics, and display graphic-text elements through a memory mapping graphic processing unit (GPU).

[0017] The basic elements of the vector graphics include points, nodes, straight lines, arc segments, circles, ellipses, multi-segment lines, curves, surfaces, bodies, composites, and point symbols, line symbols and fill symbols formed after setting symbols of normal elements.

[0018] The basic elements of the raster graphics include normal pictures and image pictures, the normal pictures include PNG, BMP, JPEG and TIFF, and the image pictures are normal pictures stored through a pyramid.

[0019] Optionally, the graphic-text interactive editing module is configured to update graphic-text elements in the intelligent graphic-text integrated processing system through parameter calling caused by receiving external input.

[0020] The editing capability of the graphic-text interactive editing module includes creation, modification and deletion of elements.

[0021] The multimedia data includes text, vector graphics, raster graphics, animations, audio and video.

[0022] Optionally, the graphic-text linkage module is configured to update associated graphic primitives based on built-in attributes of graphic-text elements, including associated calling updating through a built-in unique identifier, linkage updating through support of inheritance or combination of graphic primitive styles, and active uniform updating after matching one or more attributes.

[0023] The updating of the image-text linkage module supports automatic generation based on the image-text large model, including: through the analysis and understanding of the image-text elements and their context by the image-text large model, automatically generating and updating other multimedia elements associated with them, including graphics or text; inputting a natural language description of the modification of the local context, and automatically updating the context-related multimedia elements including graphics or text through the intent recognition and layout operation of the image-text large model.

[0024] Optionally, the construction method of the image-text large model comprises:

[0025] S1: Collecting the image-text data and the engineering field document data and preprocessing to obtain the multimedia data, the content parameterized description of the multimedia data, which includes: content outline, image-text content parameters and layout parameter definition, correlation, parameterized processing command;

[0026] S2: Extracting and understanding the features of the parameterized processing command, associating and aligning the parameterized graphic data and the parameterized text description, and establishing a mapping relationship between them;

[0027] S3: Based on the mapping relationship, an image-text large model oriented to parameterized description is constructed, a large model framework is established combining Transformer and graph neural network, pre-training is performed using parameterized data and vector graphics, raster graphics data respectively, and the pre-trained model is fine-tuned to obtain an optimized image-text large model, which accurately understands and executes the parameterized processing command, and the parameterized data includes: the image-text content parameters and the layout parameters;

[0028] S4: Based on the optimized image-text large model, the image-text large model service is provided, and the natural language input of the user is recognized to generate, update the corresponding text content or multimedia content including vector and raster graphics or layout, or generate, update the corresponding text content according to the multimedia content including vector graphics and raster graphics.

[0029] Optionally, the manual custom way of image-text processing and layout is realized by actively calling the image-text interactive editing module and passively triggering the image-text linkage module;

[0030] The automatic way of image-text processing and layout is realized by the image-text large model module, taking the template as input or intelligent analysis to automatically generate graphics, images and realize mixed layout; the intelligent image-text integrated processing system also establishes a new generation of knowledge base and expert base with intelligent deduction function based on the image-text large model, relying on pre-set graphics or other multimedia data, constructs application scene simulation of each professional full process, industry report standard verification, expert review function.

[0031] Optionally, the GIS spatial relationship and the CAD relationship each comprise a topological spatial relationship, a metric spatial relationship and an order spatial relationship.

[0032] The key attribute data further comprises built-in professional attribute data identifiable by the large graphic-text model;

[0033] The semantic analysis and correlation analysis comprises automatic labeling and topological verification.

[0034] The spatio-temporal reasoning function comprises historical data backtracking and trend prediction.

[0035] The automatically generated professional analysis report of the graphic-text comprises structured parameters generated by the large graphic-text model analyzing specific graphic-text content in the intelligent graphic-text integrated processing system, and combines graphic-text professional attributes and the GIS spatial relationship and the CAD relationship, calls built-in services in the system to generate content and calls the graphic-text layout module to complete layout, and generates an engineering special report in conformity with a preset document template format.

[0036] The built-in services comprise vector map services, raster map services, query services and extensible professional services.

[0037] The vector map services comprise creation, opening, content addition, deletion, modification and query.

[0038] The raster map services comprise cutting, shortest path and disaster avoidance route.

[0039] The query service refers to efficient query based on spatial indexing of spatial features.

[0040] The extensible professional services comprise surveying and mapping and ecological environment, resource management special map generation, pie chart generation, column chart generation, circuit diagram generation, building and mechanical equipment design diagram generation, mine special map generation, road and bridge special map generation, water conservancy special map generation, and custom professional vector graphic processing service components written by personnel according to their own engineering professional needs.

[0041] Optionally, the large graphic-text model module is further configured to record the writing style of the user and continuously iteratively upgrade, so as to ensure that the generated professional report style is consistent with the current writing style of the user.

[0042] The intelligent graphic-text integrated processing method provided by the embodiment of the application is applied to the intelligent graphic-text integrated processing system.

[0043] Through the vector graphic processing platform in the intelligent graphic-text integrated processing system, editing and processing of basic text, vector graphics, raster graphics and other multimedia data are performed, and manual custom graphic-text processing and layout are provided.

[0044] Through the generation service of the image-text large model in the intelligent image-text integrated processing system, for the scene with graphical display, the script of graphical drawing or visualization is automatically generated in combination with the context, and the corresponding vector graphics or raster graphics are generated by the intelligent image-text integrated processing system after execution; or in combination with the vector graphics or raster graphics, the corresponding text content is automatically generated;

[0045] Through the layout service of the image-text large model in the intelligent image-text integrated processing system, in combination with the specified theme or the template, the layout position or adjustment position script of the document content is automatically generated, and the layout is completed by the intelligent image-text integrated processing system after execution, forming a professional document report with image-text.

[0046] The intelligent image-text integrated processing system provided by the application comprises:

[0047] The professional vector graphics processing platform is taken as a base, unified mixed arrangement of multimedia data such as text, pictures and tables is provided, and the unified processing of image-text can be actually completed on one platform. Meanwhile, in combination with the intelligent processing requirement of image-text processing, the image-text large model with professional analysis capability is used to conveniently complete the image-text conversion, the generation and printing output of the special topic document in the intelligent image-text integrated processing system, and the intelligent support is provided for the engineering image-text processing and report writing and other applications. The intelligent image-text integrated processing system supports intelligent automatic image-text generation and layout, is especially suitable for the engineering field requiring a large amount of image-text data and having an association relationship between image and text, replaces the traditional text type office software, has a wide application prospect and high practicability. BRIEF DESCRIPTION OF DRAWINGS

[0048] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments with reference made to the accompanying drawings. The drawings are for purposes of illustration only and are not intended to be limiting in any respect. Further, like reference numerals are used throughout the drawings and textual to designate identical or like components. In the drawings:

[0049] Figure 1 is an architecture diagram of an intelligent image-text integrated processing system according to an embodiment of the application. DETAILED DESCRIPTION

[0050] In order to make the above objectives, characteristics and advantages of the application more apparent and easy to understand, the application will be further described in detail below with reference to the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the application, are only a part of the embodiments of the application, and are not used to limit the application.

[0051] At present, when processing vector or raster graphics, the traditional graphic-text processing software usually needs to be completed through external program calling (such as Microsoft OLE technology) or screenshot in the text processing system (such as WPS). This method is cumbersome to operate, needs to operate multiple software, and after data sharing across terminals and systems, it may not be able to normally display or process external graphics due to the absence of external calling programs, which has become one of the obstacles affecting the application of graphic-text systems in the Internet era.

[0052] For example: the user needs to make a power supply design report, and makes a power supply diagram through AutoCAD, and makes the final report through WPS. The report needs to insert the power supply diagram content (the data exists in the form of pictures or OLE objects) at multiple places, and give corresponding content with text description or analysis report. These descriptions or reports are usually related to the power supply diagram content, if the power supply diagram changes, the user can only re-screenshot and copy and paste it into WPS, if the attribute of the power supply diagram content changes (for example, the load of a certain device changes), the corresponding text description or analysis report in the report also needs to be modified, the overall process is relatively cumbersome to operate.

[0053] In view of the above problems, the inventors creatively propose an intelligent graphic-text integrated processing system and an intelligent graphic-text integrated processing method. The technical solutions of the present application are explained and described in detail as follows.

[0054] The intelligent graphic-text integrated processing system proposed in the present application, referring to the schematic diagram shown in Figure 1 , includes a vector graphics processing platform, a graphic-text large model module, a graphic-text layout module, and a graphic-text document generation module.

[0055] The vector graphics processing platform internally includes a multimedia data processing engine including text, vector graphics, raster graphics, animation, and video, and is used to directly complete graphic editing, text processing, and mixed layout of multimedia data with a unified graphic interaction interface. The vector graphics processing platform can not only process data of GIS and CAD platforms (including GIS graphics and CAD graphics), but also process multimedia data including text, vector graphics, raster graphics, animation, and video.

[0056] The graphic-text large model module is used to construct a graphic-text large model with intelligent understanding, processing, and layout capabilities for multimedia data by taking general graphic-text resources and engineering field document resources as data sets, combining multimedia data including vector graphics, parameterized description of multimedia data content, and layout association relationship therebetween.

[0057] The graphic-text layout module is used for providing graphic-text processing and layout in a manual customizing manner based on a vector graphics processing platform through graphic processing and interaction, and is used for providing graphic-text processing and layout in an automatic manner based on a graphic-text large model module, automatically generating vector graphics or raster graphics and layout corresponding to content based on text content according to a template or intelligent analysis, or automatically generating text content and layout corresponding to vector graphics or raster graphics based on vector graphics or raster graphics, wherein the template includes a template provided by the graphic-text large model module or provided by a third party.

[0058] The graphic-text document generation module is used for automatically generating a professional document report with pictures and text based on the graphic-text layout module, combining a given theme or template, through semantic analysis and correlation analysis of key attribute data and GIS spatial relationship data and CAD spatial relationship data, combining the graphic-text large model module and a space-time reasoning function.

[0059] In an embodiment of the present application, the vector graphics processing platform includes a graphic-text rendering module, a graphic-text interactive editing module and a graphic-text linkage module, and functions of the vector graphics processing platform are supported and realized through the modules.

[0060] Through the graphic-text rendering module, the graphic-text interactive editing module and the graphic-text linkage module, the vector graphics processing platform has editing capability of vector and raster graphics, and simultaneously provides arrangement and processing capability of multimedia data on a graphics processing base.

[0061] Through the graphic-text rendering module, the graphic-text interactive editing module and the graphic-text linkage module, a unified interactive interface is provided for the intelligent graphic-text integrated processing system, vector and raster graphics editing, text processing and mixed layout can be completed without calling external programs, seamless switching, mutual assistance, eliminating a system-level data exchange link, and avoiding loss of shielding information including attribute and topological relationship.

[0062] Among them, the mixed layout refers to unified operation of vector and raster graphics and multimedia data, supporting tiled sequential layout of office software through a global offset attribute built in a graphic element, or supporting stacked placeholder layout through a fixed position attribute of the graphic element.

[0063] The unified interactive interface without calling external programs refers to rendering, editing layout and linkage updating of multimedia data including vector and raster graphics, which are respectively completed by the graphic-text rendering module, the graphic-text interactive editing module and the graphic-text linkage module without external programs; a rendering area of the unified interactive interface is dynamically updated with graphic elements through preset or dynamic calculation of a critical value, has stepless zooming capability, and supports browsing and interactive editing in multiple view granularities.

[0064] In an embodiment of the present application, the graphic-text rendering module is used to load vector and raster graphics and complete display of graphic-text elements through a memory-mapped graphics processing unit (GPU); the basic elements of the vector graphics include points, nodes, straight lines, arc segments, circles, ellipses, polyline, curves, surfaces, solids, composites, and point symbols, line symbols, and fill symbols formed after setting symbols of common elements; the basic elements of the raster graphics include common pictures and image pictures, the common pictures include PNG, BMP, JPEG, and TIFF, and the image pictures are common pictures stored through a pyramid.

[0065] Preferably, the graphic-text interactive editing module can update the graphic-text elements in the intelligent graphic-text integrated processing system through parameter calling triggered by receiving external input; the editing capabilities of the graphic-text interactive editing module include creation, modification, and deletion of elements; the multimedia data includes text, vector graphics, raster graphics, animations, audio, and video.

[0066] In an embodiment of the present application, the graphic-text linkage module is used to complete update of associated graphic elements based on built-in attributes of graphic-text elements, including: update through associated calling of built-in unique identifiers, update through graphic element style linkage supporting inheritance or combination, and active uniform update after matching one or more attributes.

[0067] The update of the graphic-text linkage module supports automatic generation based on a graphic-text large model, including: automatic generation and update of other multimedia elements including graphics or text associated with the graphic-text elements through analysis and understanding of contexts of the graphic-text elements by the graphic-text large model; automatic update of multimedia elements related to contexts including graphics or text through intention recognition and layout operation of the graphic-text large model based on natural language input of modification description of local contexts.

[0068] In an embodiment of the present application, a construction method of the graphic-text large model includes:

[0069] S1: collect graphic-text data and engineering field document data and perform preprocessing to obtain multimedia data including vector graphics and raster graphics, content parameterized description of the multimedia data, which includes: content outline, graphic-text content parameters, and definition of layout parameters, association relationship, and parameterized processing commands;

[0070] S2: extract features and understand semantics of the parameterized processing commands, associate and align the parameterized graphic data and the parameterized text description, and establish a mapping relationship therebetween;

[0071] S3: Based on the mapping relationship, a graph-text large model oriented to parameterized description is constructed, a large model framework is established combining Transformer and graph neural network, pre-training is performed using parameterized data and vector graphics and raster graphics respectively, and the pre-trained model is fine-tuned to obtain an optimized graph-text large model, so that the parameterized processing command can be accurately understood and executed, wherein the parameterized data includes the graph-text content parameters and the layout parameters in step S1;

[0072] S4: Based on the optimized graph-text large model, a service of the graph-text large model is provided, and a natural language input of a user is recognized to generate or update corresponding text content or multimedia content including vector and raster graphics or layout, or to generate or update corresponding text content according to the multimedia content including vector graphics and raster graphics.

[0073] In an embodiment of the present application, the manual customization mode of graph processing and layout is realized by actively calling the graph-text interactive editing module and passively triggering the graph-text linkage module. For example:

[0074] When the user is making a ventilation system diagram, the user creates a ventilation node object, a ventilation branch object and a text or multi-line text object through the graph-text interactive editing module in the vector graphics processing platform. The ventilation node and the ventilation branch are professional objects, which include professional attribute information, associated object information and control point information. These objects are created through mouse interaction operation. After the user selects a professional object, the control point information of the professional object is displayed through the rendering module, and the attribute information of the professional object is displayed through the attribute dialog box of the interactive editing module. The user can move the position of the professional object or perform other geometric transformations by dragging the control point, and can modify the attributes of the professional object by setting the attribute values in the attribute box. The geometry of the object can also change when the attribute is modified. The user can directly edit the text or multi-line text by double-clicking the object in the what-you-see-is-what-you-get editor. The graph-text interactive editing module completes the integrated editing of vector graphics and text without calling any external program. When the position of a ventilation node changes, the graph-text linkage module searches for the associated ventilation branch object in the loaded ID-ID mapping table, and calls the ventilation professional algorithm through a virtual interface to update the position of the ventilation branch object. The number information of the ventilation node can be stored in the form of a field in the text or multi-line text. When the ventilation system diagram is updated and the node number changes, the graph-text linkage module searches for the text or multi-line text by traversing the loaded field-ID mapping table and updates the field content. The graph-text association module passively triggers the update of the position and content of the vector data and the text, and then completes the layout.

[0075] The automatic processing and layout of the text and graphics is automatically generated by the text and graphics large model module, and the text and graphics large model module is used as input or intelligent analysis of a template (provided by the text and graphics large model module or provided by a third party) to automatically generate graphics, images and realize mixed layout. Based on the text and graphics large model, the intelligent text and graphics integrated processing system establishes a new generation of knowledge base and expert base with intelligent deduction function, constructs various professional full-process application scene simulation, industry report standard verification, expert review and other functions by relying on the pre-set graphics or other multimedia data.

[0076] The GIS spatial relationship and the CAD spatial relationship involved in the above description include: topological spatial relationship, metric spatial relationship and sequential spatial relationship. In addition to the spatial relationship, the key attribute data also includes built-in professional attribute data that can be recognized by the text and graphics large model.

[0077] The semantic analysis and correlation analysis include: automatic labeling and topological verification. The spatio-temporal reasoning function includes: historical data backtracking and trend prediction.

[0078] The automatically generated professional analysis report of the text and graphics refers to: the structured parameters generated by the intelligent text and graphics integrated processing system through the analysis of the text and graphics large model, and the combination of the text and graphics spatial relationship and the professional attribute, and the built-in service of the system, and the generation of the professional analysis report in accordance with the preset document template format.

[0079] The built-in service includes: vector map service, raster map service, query service and extensible professional service. The vector map service includes: creation, opening, content addition, deletion, modification and query. The raster map service includes: cutting, shortest path and disaster avoidance route. The query service refers to the efficient query based on the spatial index of the spatial feature. The extensible professional service includes: surveying and mapping and ecological environment, resource management thematic map generation, pie chart generation, column chart generation, circuit diagram generation, building and mechanical equipment design diagram generation, mine thematic map generation, road and bridge thematic map generation, water conservancy thematic map generation, and custom professional vector graphics processing service components written by personnel according to their own engineering professional needs.

[0080] In an embodiment of the present application, the text and graphics large model module is also used to record the writing style of the user and continuously iterate and upgrade to ensure that the generated professional report style is consistent with the current writing style of the user.

[0081] Based on the above-mentioned intelligent text and graphics integrated processing system, the present application further provides an intelligent text and graphics integrated processing method, which is applied to the above-mentioned intelligent text and graphics integrated processing system, and the method comprises:

[0082] Step T1: Through the vector graphics processing platform in the intelligent image-text integrated processing system, the editing and processing of basic text, vector graphics, raster graphics and other multimedia data are provided with manual customization mode of image-text processing and layout;

[0083] Step T2: Through the generation service of the image-text large model in the intelligent image-text integrated processing system, for the scene with graphical display, the script of graphical drawing or visualization is automatically generated in combination with the context, and the corresponding vector graphics or raster graphics are generated after execution by the intelligent image-text integrated processing system; or in combination with the vector graphics or raster graphics, the corresponding text content is automatically generated.

[0084] Step T3: Through the layout service of the image-text large model in the intelligent image-text integrated processing system, the layout position or adjustment position script of the document content is automatically generated in combination with the specified theme or template, and the layout is completed after execution by the intelligent image-text integrated processing system, forming a professional document report with image-text and user writing style.

[0085] In order to better understand the intelligent image-text integrated processing system and method proposed in the present application, the intelligent image-text integrated processing system is described in detail how to layout and how to automatically generate a professional analysis report in combination with the foregoing various modules.

[0086] For example: a user makes a professional power supply design report through the intelligent image-text integrated processing system proposed in the present application.

[0087] Manual layout:

[0088] (1) The user inputs an existing power supply system diagram or generates an instant power supply system diagram in the intelligent image-text integrated processing system, which usually contains:

[0089] ① Power supply equipment, power supply cable mechanical and electrical professional objects, which are displayed as vector graphics complexes in the system, and the complex is composed of points, straight line segments, text, etc.; the object contains professional attributes (such as voltage, current, etc.); the object contains spatial correlation information, for example, the power supply equipment contains associated cable data information, and the power supply cable also stores associated power supply equipment information; the object contains control point information, for example, the power supply equipment contains a center point control point for translation;

[0090] 2) Legend information, which is displayed as a raster graphic in the system; the vector graphic processing platform unifies the vector graphics, raster graphics and text in the system diagram in one system through a graphic-text rendering module to complete rendering, without relying on any third-party software. The vector graphics are rendered in the GPU through point elements, line elements and triangle elements, the raster graphics are rendered in the GPU through texture elements, the text is rendered in the GPU through triangle elements or texture elements, and the control points are rendered in the GPU as triangle elements.

[0091] (2) The user clicks to insert text or multi-line text in the vector graphic at the designated position of the system diagram to form the text content of the report. The text and multi-line text both contain text format information, and the multi-line text also contains paragraph information for setting the paragraph format. The text or multi-line text interactive editor has a what-you-see-is-what-you-get visual effect. After the text or multi-line text input is completed, it can be re-entered into the editing state for content modification by double-clicking. After the vector graphic in the system diagram is selected, the rendering module displays its control point information. The user can edit the vector graphic as a whole or in part by selecting and dragging, for example, after the central control point of the power supply device is selected, the user can complete the translation of the entire power supply device object by dragging. For another example, after a control point of the ordinary polyline in the graphic is selected, the user can move the corresponding point of the control point by dragging. After a professional object in the intelligent graphic-text integrated processing system diagram is selected, the interactive editing module displays the attribute information of the object through an attribute dialog box. The user edits the attribute content in the attribute dialog box to modify the attribute. For example, the user modifies the length of the power cable object through the attribute dialog box. The attribute of the object is modified, and the length of the cable in the corresponding complex body is also changed, finally notifying the rendering module to refresh and display the new text. Thus, the vector graphic processing platform completes the mixed editing of the vector graphic and the text through the unified graphic-text interactive editing module.

[0092] (3) In the power supply system diagram, when the user moves a power supply device, the associated nodes of the power supply cable associated with it need to change accordingly. The association between objects is stored in the figure-text association module in the form of ID-ID mapping, and this association is formed when the power supply device-power supply cable is created and associated when the system diagram is drawn or generated. When the user moves the power supply device, the figure-text association module reads the ID of the object moved by the user, queries the associated power supply cable ID through mapping, and then calls the professional update algorithm through the virtual interface to complete the update of the power supply cable object. In the power supply system diagram, the user inserts the length of a certain power supply cable in a paragraph through multiple text fields. The association between the field and the object is stored in the figure-text association module in the form of field type-ID mapping, and this association is formed when the field is created. When the user sets the length of the power supply cable, the figure-text association module traverses the field types in the system diagram and updates the field value in the multi-line text after matching the ID of the power supply cable. Thus, the vector graphics processing platform completes the associated meta update based on the built-in properties of the figure-text elements through the figure-text association module.

[0093] (4) When the user needs to fine-tune the power supply cable in the power supply system diagram, the view can be scaled to a view scale that is convenient for the user to interact and edit.

[0094] Automatic layout:

[0095] (1) Collect general power industry figure-text resources and document resources as data sets to obtain power supply type graphic data and multimedia data and their layout association. Parameterize the power supply device and power supply cable information in the graphic data, for example, in the form of Json string. On this basis, train and build a figure-text large model with intelligent understanding, processing, and layout of figure-text and multimedia.

[0096] (2) The user transmits the text description information of the power supply system diagram to the diagram-text large model module in the diagram-text processing system in a natural language manner. The diagram-text large model module calculates the Json string description information of the power supply system diagram according to the training results. The diagram-text media processing engine creates objects, adds attributes, and establishes associated information according to the professional objects or text, multi-line text, and raster pictures in the Json string. The professional objects include mechanical and electrical equipment objects and mechanical and electrical cable objects. The text or multi-line text includes report content. The raster pictures include legend information in the system diagram. The diagram-text layout module automatically generates vector graphics or raster graphics corresponding to the text content and further automatically lays out all the created objects according to the layout information in the Json string. The user can upload a self-defined format power supply design report template as needed. The diagram-text layout module can also generate content and layout according to this template.

[0097] Automatic generation of power supply design report:

[0098] In automatic layout, the user uses natural language as a parameter to call vector graphics or raster graphics and complete layout through the diagram-text large model module. In order to further generate a professional power supply design report with text and pictures, the following operations are needed:

[0099] (1) Use the attribute information generated by automatic layout for measurement analysis. For example, use the built-in mechanical and electrical service algorithm to count the setting value information of all high and low voltage switches, obtain the tree structure diagram of the power supply model, and return the system in SVG format. At the same time, the mechanical and electrical service algorithm analyzes the coordination relationship of each switch, makes a preliminary judgment on "overstep tripping", and makes a decision on the value of switch setting value, and returns the system in multi-line text;

[0100] (2) Use the associated information generated by automatic layout for power supply equipment and cable topology analysis. Use the associated information of all power supply equipment and power supply cables as parameters to call the built-in mechanical and electrical service algorithm to calculate the topology structure information of the power supply system diagram, and return the system in multi-line text;

[0101] (3) In the power supply design report template uploaded by the user or the diagram-text large model generation template, replace the preset fields with the attribute information generated by automatic layout; and calculate the results in the template by calling the built-in mechanical and electrical service algorithm. The results are returned to the system in SVG format after being formed by the built-in cutout service;

[0102] (4) In the power supply design report template uploaded by the user or the diagram-text large model generation template, for the preset graphic content, generate SVG or Json format and return the system for filling by calling the built-in local diagramming service; for the professional analysis and description of the graphic content, get the multi-line text returned by the system by calling the built-in mechanical and electrical service algorithm;

[0103] (5) After the user gets the automatically generated power supply design report, manual fine-tuning can be performed again.

[0104] (6) The formed power supply design report is circulated into the picture-text large model for iterative training and upgrading, so that the picture-text integrated processing system gradually has the ability to generate a report in the user's writing style.

[0105] The intelligent picture-text integrated processing system provided by the application takes a professional graphics processing platform as a base, provides unified mixed arrangement of multimedia data such as text, pictures and tables, and can truly complete the unified processing of pictures and texts on one platform. Meanwhile, combined with the intelligent processing requirements of picture-text processing, based on a picture-text large model with professional analysis capability, the intelligent picture-text integrated processing system can conveniently complete picture-text mutual conversion, generation of special topic documents and printing output, and provides intelligent support for engineering picture-text processing and report writing and the like. The intelligent picture-text integrated processing system supports intelligent automatic picture-text generation and layout, is especially suitable for fields such as engineering that need a large amount of picture-text data and have associated relationships between pictures and texts, replaces traditional text-type office software, has a broad application prospect and high practicability.

[0106] Although the preferred embodiments of the embodiments of the application have been described, those skilled in the art can make further changes and modifications to the embodiments once they know the basic inventive concept. Therefore, the appended claims are intended to cover the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the application.

[0107] Finally, it should be noted that, in this document, the terms such as first and second are used merely to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprising", "containing" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or terminal device. Without more limitations, the element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, article or terminal device including the element.

[0108] The embodiments of the present application are described above with reference to the accompanying drawings, but the present application is not limited to the above-described specific embodiments, and the above-described specific embodiments are merely illustrative, but not restrictive, and a person of ordinary skill in the art can make many forms under the inspiration of the present application without departing from the purpose of the present application and the scope protected by the claims, and these all belong to the protection of the present application.

Claims

1. An intelligent integrated processing system of text and graphics, characterized in that, Include: Vector graphics processing platform, picture-text large model module, picture-text layout module, picture-text document generation module; The vector graphics processing platform, built-in multimedia data processing engine including text, vector graphics, raster graphics, animation, video, for with a unified graphical interface, directly complete graphics editing, text processing and mixed layout of multimedia data; The picture-text large model module is used for taking general picture-text resources and engineering field document resources as data sets, combining multimedia data including text, vector graphics, raster graphics, animation, video, parameterization description of multimedia data content and layout association relationship among them, and constructing a picture-text large model with intelligent understanding, processing and layout capability of multimedia data; The picture-text layout module is used for based on the vector graphics processing platform, providing manual customization method of picture-text processing and layout through graphical processing and interaction; And, for based on the picture-text large model module, provide automatic way of picture-text processing and layout, according to template or intelligent analysis, based on text content to automatically generate corresponding vector graphics or raster graphics and layout, or based on vector graphics or raster graphics to automatically generate corresponding text content and layout, the template includes: the picture-text large model module provides or third party provides; The picture-text document generation module is used for based on the picture-text layout module, combining given theme or the template, through semantic analysis and correlation analysis key attribute data and GIS spatial relationship data, CAD spatial relationship data, combining picture-text large model module and space-time reasoning function, automatically generate picture-text professional document report, the GIS spatial relationship data, the CAD spatial relationship data all include topological relationship data; Wherein, the vector graphics processing platform includes: picture-text rendering module, picture-text interactive editing module and picture-text linkage module; through the picture-text rendering module, picture-text interactive editing module and picture-text linkage module, so that the vector graphics processing platform has the editing capability of vector graphics and raster graphics, and simultaneously provides the arrangement processing capability of multimedia data on the basis or platform of graphics processing. The intelligent image-text integrated processing system is provided with a unified interactive interface by the image-text rendering module, the image-text interactive editing module and the image-text linkage module, and vector and raster graphics editing, text processing and mixed layout can be completed without calling external programs, seamless switching, mutual assistance, eliminating system-level data exchange links, avoiding the loss of key information including attributes and topological relations; the mixed layout refers to unified operation on vector graphics, raster graphics and the multimedia data, supporting the tiling sequential layout of office software through the built-in global offset attribute of the graphic element, or supporting the stacking placeholder layout through the fixed position attribute of the graphic element; the rendering, editing and layout of the multimedia data are completed by the image-text rendering module, the image-text interactive editing module and the image-text linkage module without external programs; the rendering area of the unified interactive interface dynamically updates graphic elements through a preset or dynamically calculated critical value, has stepless zooming capability and supports browsing and interactive editing under multiple view granularities; The construction method of the image-text large model comprises the following steps: S1: collecting image-text data and engineering field document data and preprocessing to obtain multimedia data and content parameterized description of the multimedia data, which comprises content outline, image-text content parameters and layout parameter definition, correlation, parameterized processing command; S2: extracting and understanding the features of the parameterized processing command, associating and aligning the parameterized graphic data and parameterized text description, and establishing the mapping relationship therebetween; S3: based on the mapping relationship, constructing an image-text large model oriented to parameterized description, establishing a large model framework combining Transformer and graph neural network, pre-training using parameterized data and vector graphics and raster graphics respectively, and fine-tuning the pre-trained model to obtain an optimized image-text large model, which accurately understands and executes the parameterized processing command; the parameterized data comprises the image-text content parameters and the layout parameters; S4: based on the optimized image-text large model, providing image-text large model services and identifying user natural language input to generate or update corresponding text content or multimedia content including vector and raster graphics or layout, or generating or updating corresponding text content according to vector graphics and raster graphics.

2. The intelligent image-text integrated processing system according to claim 1, wherein The image-text rendering module is used to load the vector and raster graphics and display image-text elements through a memory mapping graphics processing unit (GPU). The basic elements of vector graphics include points, nodes, straight lines, arc segments, circles, ellipses, multi-segment lines, curves, surfaces, bodies, complexes, and point symbols, line symbols and fill symbols formed after setting symbols of basic elements. The basic elements of raster graphics include ordinary pictures and image pictures, the ordinary pictures include PNG, BMP, JPEG and TIFF, and the image pictures are ordinary pictures stored through the establishment of a pyramid.

3. The intelligent image-text integrated processing system of claim 1, wherein The picture-text interactive editing module updates the picture-text elements in the intelligent picture-text integrated processing system through parameter calling triggered by receiving external input; The editing capability of the picture-text interactive editing module includes creation, modification and deletion of elements; The multimedia data includes text, vector graphics, raster graphics, animation, audio and video.

4. The intelligent image-text integrated processing system according to claim 1, wherein The picture-text linkage module updates the associated graphics based on the built-in attributes of the picture-text elements, including: updating through built-in unique identification association, updating through graphics style linkage supporting inheritance or combination, and actively updating after one or more attribute matching; The picture-text linkage module update supports automatic generation based on the picture-text large model, including: automatically generating and updating other multimedia elements including graphics or text associated with the picture-text elements through the picture-text large model analysis and understanding of the context; modifying the local context in natural language mode, and automatically updating the context-related multimedia elements including graphics or text through the intent recognition and layout operation of the picture-text large model.

5. The intelligent image-text integrated processing system according to claim 1, wherein The picture-text processing and layout in the manual customization mode are achieved by actively calling the picture-text interactive editing module and passively triggering the picture-text linkage module; The automatic picture-text processing and layout is achieved by the picture-text large model module, which generates graphics, images and realizes mixed layout based on the template as input or intelligent analysis; the intelligent picture-text integrated processing system also establishes a new generation of knowledge base and expert base with intelligent deduction function based on the picture-text large model and preloaded graphics or other multimedia data, and constructs application scene simulation, industry report standard verification and expert review function for each professional full process.

6. The intelligent image-text integrated processing system according to claim 1, wherein The GIS spatial relationship and the CAD spatial relationship both include topological spatial relationship, metric spatial relationship and sequential spatial relationship; The key attribute data also includes built-in professional attribute data that can be recognized by the picture-text large model; The semantic analysis and association analysis include automatic labeling and topological verification; The spatio-temporal reasoning function includes historical data backtracking and trend prediction; The automatically generated picture-text professional analysis report refers to the structured parameters generated by the picture-text large model analyzing the picture-text content in the intelligent picture-text integrated processing system, combined with picture-text professional attributes and GIS space and CAD space relationship, calling system built-in services to generate content and calling the picture-text layout module to complete layout, generating engineering topic report in accordance with the preset document template format; The built-in services include vector map service, raster map service, query service and extensible professional service; The vector map service includes creation, opening, content addition, deletion, modification and query; The raster map service includes cutting, shortest path and disaster avoidance route; The query service refers to efficient query based on spatial index of spatial features; The extensible professional service includes: surveying and mapping and ecological environment, resource management thematic map generation, pie chart generation, column chart generation, circuit diagram generation, building and mechanical equipment design diagram generation, mine thematic map generation, road and bridge thematic map generation, water conservancy thematic map generation, and custom professional vector graph processing service components written by personnel according to their own engineering professional needs.

7. The intelligent image-text integrated processing system of claim 1, wherein The large model module is also used for recording the writing style of the user and continuously iterating and upgrading to ensure that the generated professional report style is consistent with the current writing style of the user.

8. An intelligent image-text integration processing method, characterized in that, The intelligent graphics-text integrated processing method is applied to the intelligent graphics-text integrated processing system in any of claims 1-7, and comprises: Through the vector graph processing platform in the intelligent graphics-text integrated processing system, the editing and processing of basic text, vector graphics, raster graphics and other multimedia data are provided with manual customization mode of graphics-text processing and typesetting; wherein the vector graph processing platform comprises a graphics-text rendering module, a graphics-text interactive editing module and a graphics-text linkage module; through the graphics-text rendering module, the graphics-text interactive editing module and the graphics-text linkage module, the vector graph processing platform has the editing capability of vector graphics and raster graphics, and simultaneously provides the arrangement processing capability of multimedia data on the graphics processing base or platform; Through the graphics-text rendering module, the graphics-text interactive editing module and the graphics-text linkage module, a unified interactive interface is provided for the intelligent graphics-text integrated processing system, without calling external programs to complete vector and raster graphics editing, text processing and mixed typesetting, seamless switching, mutual assistance, eliminating system-level data exchange links, avoiding the loss of key information including attributes and topological relationships; wherein the mixed typesetting refers to unified operation of vector graphics, raster graphics and the multimedia data, supporting the tiling sequential typesetting of office software through the global offset attribute of the built-in graphics element, or supporting the stacking placeholder typesetting through the fixed position attribute of the graphics element; the unified interactive interface, without calling external programs, refers to that the rendering, editing and typesetting and linkage updating of the multimedia data are respectively completed by the graphics-text rendering module, the graphics-text interactive editing module and the graphics-text linkage module without the aid of external programs; the rendering area of the unified interactive interface dynamically updates the graphic elements through the preset or dynamically calculated critical value, has stepless zooming capability, and supports browsing and interactive editing under multiple view granularities; Through the generation service of the graphic-text large model in the intelligent graphic-text integrated processing system, for the scene with graphical display, the script of graphical drawing or visualization is automatically generated in combination with the context, and the corresponding vector graphics or raster graphics are generated after being executed by the intelligent graphic-text integrated processing system; or in combination with the vector graphics or raster graphics, the corresponding text content is automatically generated; the construction method of the graphic-text large model comprises: S1: collecting the graphic-text data and the engineering field document data and pre-processing to obtain the multimedia data, content parameterized description of the multimedia data, which comprises: content outline, graphic-text content parameters and layout parameter definition, correlation, parameterized processing command; S2: extracting and understanding the features of the parameterized processing command, associating and aligning the parameterized graphic data and parameterized text description, and establishing the mapping relationship therebetween; S3: based on the mapping relationship, a graphic-text large model facing parameterized description is constructed, a large model framework is established in combination with Transformer and a graph neural network, parameterized data and vector graphics, raster graphics are respectively pre-trained, and the pre-trained model is fine-tuned to obtain an optimized graphic-text large model, so that the graphic-text large model accurately understands and executes the parameterized processing command; the parameterized data comprises the graphic-text content parameters and the layout parameters; S4: based on the optimized graphic-text large model, a service of the graphic-text large model is provided, and the natural language input of a user is recognized to generate or update the corresponding text content or multimedia content including vector and raster graphics or layout, or to generate or update the corresponding text content according to the multimedia content including vector graphics and raster graphics; Through the layout service of the graphic-text large model in the intelligent graphic-text integrated processing system, in combination with a specified theme or the template, a layout position or adjustment position script of the document content is automatically generated, and after being executed by the intelligent graphic-text integrated processing system, the layout is completed to form a professional document report with graphic-text combination and user writing style.

Citation Information

Patent Citations

  • Vector data editing method, device and system

    CN107807911A

  • Content display method and device based on large language model, equipment, storage medium and program product

    CN119557435A