Image forming method, apparatus, server, storage medium, and program product

By using a large language model to generate instructions for processing target images and documents in an image forming apparatus, the problems of low intelligence and poor security are solved, and intelligent control and security are improved.

CN122113857APending Publication Date: 2026-05-29ZHUHAI PANTUM ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202512042491.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In the existing technology, the process of acquiring target images by image forming apparatus cannot be intelligently controlled and has poor security, which cannot meet the working scenarios with high security requirements.

Method used

Image acquisition instructions and document processing instructions are generated by a large language model. The image forming device performs the processing of the target image and the document to be processed locally, avoiding uploading to the server.

Benefits of technology

Intelligent control of the target image acquisition process has been achieved, which has improved the intelligence level of the image forming device and enhanced its safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122113857A_ABST
    Figure CN122113857A_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides an image forming method and device, a server, a storage medium and a program product. The method comprises: obtaining a to-be-processed document and description information, the description information comprising an image acquisition strategy and a document processing strategy; determining an image acquisition instruction and a document processing instruction generated by a large language model according to the description information; executing an operation according to the image acquisition instruction to obtain a target image; and processing the target image and the to-be-processed document according to the document processing instruction to obtain a processed document. In this embodiment, the image acquisition instruction is generated by the large language model, and then the operation of acquiring the target image can be executed according to the image acquisition instruction, thereby realizing intelligent control of the target image acquisition process and improving the intelligent degree. In addition, the target image and the to-be-processed document are processed at the image forming device end, and since the image forming device does not need to upload the acquired target image and the to-be-processed document to the server, the security can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image forming technology, and specifically to an image forming method, apparatus, server, storage medium, and program product. Background Technology

[0002] With the development of artificial intelligence technology, some image forming devices can be used in conjunction with AI tools to meet users' personalized and diverse needs. For example, users may need to use external AI tools to process the target image and the document to be processed acquired by the image forming device. This could involve inserting the target image into the document or replacing the target image with the background of the document.

[0003] To address the aforementioned user needs, one implementation scheme in the relevant technology is as follows: First, the user sends the document to be processed to the image forming apparatus via a terminal device; then, the image forming apparatus uploads the obtained target image and the document to be processed together to the server; finally, the server uses artificial intelligence tools to process the target image and the document to be processed accordingly, obtains the processed document, and sends the processed document to the user's terminal device.

[0004] The above-mentioned implementation schemes in the related technologies have at least the following disadvantages: 1) They cannot intelligently control the process of the image forming device acquiring the target image based on artificial intelligence tools, resulting in a low level of intelligence; 2) The target image and the document to be processed obtained by the image forming device need to be uploaded to the server, which cannot meet the requirements of working scenarios with high security requirements, resulting in poor security.

[0005] It should be noted that the information disclosed in the background section of this application is intended only to enhance the understanding of the general background of this application, and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention

[0006] This application provides an image forming method, apparatus, server, storage medium, and program product to address the problems of low intelligence and poor security in the prior art when using artificial intelligence tools to process target images and documents acquired by image forming apparatuses.

[0007] In a first aspect, embodiments of this application provide an image forming method applied to an image forming apparatus, the method comprising: Obtain the document to be processed and its description information, wherein the description information includes an image acquisition strategy and a document processing strategy; Based on the description information, image acquisition instructions and document processing instructions generated by the large language model are determined; wherein, the image acquisition instructions are determined according to the image acquisition strategy, and the document processing instructions are determined according to the document processing strategy; Execute the operation according to the image acquisition instruction to acquire the target image; The target image and the document to be processed are processed according to the document processing instructions to obtain the processed document.

[0008] In one possible implementation, the image acquisition instruction is a scanning instruction, and the step of performing an operation according to the image acquisition instruction to acquire the target image includes: performing a scanning operation according to the scanning instruction to acquire the target image.

[0009] In one possible implementation, determining the image acquisition instructions and document processing instructions generated by the large language model based on the description information includes: The description information is sent to the server, which processes the description information using a first language model to generate corresponding image acquisition instructions and document processing instructions. Receive the image acquisition instruction and the document processing instruction sent by the server.

[0010] In one possible implementation, the description information further includes network retrieval description information, and the server is further configured to perform a retrieval in the network based on the network retrieval description information to obtain the target image; the step of processing the target image and the document to be processed according to the document processing instructions to obtain the processed document further includes: Receive the target image sent by the server; process the target image and the document to be processed according to the document processing instructions to obtain the processed document.

[0011] In one possible implementation, the description information further includes network-attached retrieval description information, which is used to instruct the retrieval of files related to the target image in the network based on the target image. After obtaining the target image by performing an operation according to the image acquisition instruction, the method further includes: The target image is sent to the server, and the server is further configured to retrieve files related to the target image in the network based on the network additional retrieval description information to obtain the target file; The step of processing the target image and the document to be processed according to the document processing instructions to obtain the processed document includes: Receive the target file sent by the server; The target image, the target file, and the document to be processed are processed according to the document processing instructions to obtain the processed document.

[0012] In one possible implementation, the description information is sent to a server, whereby the server processes the description information using a first language model to generate corresponding image acquisition instructions and document processing instructions, including: The description information and enhancement information are sent to the server, which processes the description information and enhancement information using a first major language model to generate corresponding scanning instructions and document processing instructions. The enhanced information includes the default image acquisition strategy description information of the image forming apparatus and / or knowledge related to the description information in the knowledge base of the image forming apparatus.

[0013] In one possible implementation, determining the image acquisition instructions and document processing instructions generated by the large language model based on the description information includes: The description information is processed using a second major language model to generate the image acquisition instructions and the document processing instructions.

[0014] In one possible implementation, the step of processing the descriptive information using a second major language model to generate the image acquisition instructions and the document processing instructions includes: The second major language model is used to process the descriptive information and enhanced information to generate image acquisition instructions and document processing instructions; The enhanced information includes the default image acquisition strategy description information of the image forming apparatus and / or knowledge related to the description information in the knowledge base of the image forming apparatus.

[0015] In one possible implementation, the step of processing the document to be processed according to the document processing instructions and the target image in accordance with the document processing strategy to obtain a processed document includes: The target image is inserted into the document to be processed to obtain the processed document; Alternatively, the target image can be replaced with the background of the document to be processed to obtain the processed document; Alternatively, the document to be processed can be adjusted according to the style of the target image to obtain a processed document; Alternatively, some or all of the text can be extracted from the target image and inserted into the document to be processed to obtain the processed document.

[0016] In one possible implementation, after obtaining the processed document, the following is also included: The processed document is sent to the terminal device, and the processed document is used for preview display on the display interface of the terminal device.

[0017] One possible implementation also includes: Receive new description information sent by the terminal device, the new description information including description information of new image acquisition strategy and / or document processing strategy; Based on the new description information and the historical description information, new image acquisition instructions and new document processing instructions generated by the large language model are determined; The operation is executed according to the new image acquisition instruction to obtain a new target image; The document to be processed is processed according to the new document processing instructions and the new target image to obtain a new processed document.

[0018] In one possible implementation, after processing the document to be processed according to the new document processing instructions and the new target image to obtain a new processed document, the method further includes: The enhanced information is updated according to the new image acquisition instructions and the new document processing instructions; The enhanced information includes the default image acquisition strategy description information of the image forming apparatus and / or knowledge related to the description information in the knowledge base of the image forming apparatus.

[0019] One possible implementation also includes: Receive a query instruction sent by the server, the query instruction being used to query the functions supported by the image forming apparatus; Send the functions supported by the image forming apparatus to the server.

[0020] Secondly, embodiments of this application provide an image forming method applied to a server, the method comprising: Receive description information sent by the image forming apparatus, the description information including an image acquisition strategy and a document processing strategy; The first major language model is used to process the descriptive information to generate corresponding image acquisition instructions and document processing instructions; Send the image acquisition instruction and the document processing instruction to the image forming apparatus; The image acquisition instruction is used to instruct the image forming apparatus to perform an operation corresponding to the image acquisition strategy; the document processing instruction is used to instruct the image forming apparatus to process the target image and the document to be processed according to the document processing strategy, thereby obtaining a processed document.

[0021] In one possible implementation, the description information further includes network retrieval description information, and the method further includes: The target image is obtained by searching the network based on the network retrieval description information. The target image is sent to the image forming apparatus, and the document processing instruction is used to instruct the image forming apparatus to process the target image and the document to be processed according to the document processing strategy.

[0022] In one possible implementation, the description information further includes network-attached retrieval description information, and the method further includes: Receive the target image sent by the image forming apparatus; Based on the network additional retrieval description information, files related to the target image are retrieved in the network to obtain the target file.

[0023] In one possible implementation, receiving the description information sent by the image forming apparatus includes: receiving description information and enhancement information sent by the image forming apparatus, wherein the enhancement information includes the image acquisition strategy description information of the image forming apparatus by default and / or knowledge related to the description information in the knowledge base of the image forming apparatus; The step of processing the descriptive information using the first major language model to generate corresponding image acquisition instructions and document processing instructions includes: processing the descriptive information and the enhanced information using the first major language model to generate corresponding image acquisition instructions and document processing instructions.

[0024] In one possible implementation, sending the image acquisition instruction and the document processing instruction to the image forming apparatus includes: Send a query command to the image forming apparatus, the query command being used to query the functions supported by the image forming apparatus; Receive the functions supported by the image forming apparatus sent by the image forming apparatus; If the functions supported by the image forming apparatus match the image acquisition instruction and the document processing instruction, then the image acquisition instruction and the document processing instruction are sent to the image forming apparatus.

[0025] One possible implementation also includes: Receive new description information sent by the image forming apparatus, the new description information including description information of new image acquisition strategy and / or document processing strategy; The new descriptive information is processed using the first major language model to generate new image acquisition instructions and new document processing instructions; Send the new image acquisition instruction and the new document processing instruction to the image forming apparatus; The image acquisition instruction is used to instruct the image forming apparatus to perform a corresponding scanning operation; the document processing instruction is used to instruct the image forming apparatus to perform corresponding processing on the document to be processed based on the target image.

[0026] Thirdly, embodiments of this application provide an image forming apparatus, including: A controller configured to perform the method described in any one of the first aspects.

[0027] Fourthly, embodiments of this application provide a server, including: processor; Memory; And a computer program, wherein the computer program is stored in the memory, and when the computer program is executed by the processor, causes the server to perform the method described in any one of the second aspects.

[0028] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in any one of the first and second aspects.

[0029] Sixthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method described in any one of the first and second aspects.

[0030] Compared with the prior art, the technical solution provided in this application has at least the following advantages: 1) Image acquisition instructions are generated by a large language model, and then the operation of acquiring the target image can be executed according to the image acquisition instructions, so as to realize intelligent control of the target image acquisition process and improve the level of intelligence; 2) The target image and the document to be processed are processed at the image forming apparatus. Since the image forming apparatus does not need to upload the acquired target image and the document to be processed to the server, security can be improved. Attached Figure Description

[0031] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1This is a schematic diagram of an application scenario provided by an embodiment of this application; Figure 2 A schematic flowchart of an image forming method provided in an embodiment of this application; Figure 3 This is a schematic diagram illustrating another application scenario provided by an embodiment of this application; Figure 4A and 4B This is a schematic diagram illustrating another application scenario provided by an embodiment of this application; Figure 5A and 5B This is a schematic diagram illustrating another application scenario provided by an embodiment of this application; Figure 6A and 6B This is a schematic diagram illustrating another application scenario provided by an embodiment of this application; Figure 6C and 6D This is a schematic diagram illustrating another application scenario provided by an embodiment of this application; Figure 7 A schematic flowchart illustrating another image forming method provided in an embodiment of this application; Figure 8A and 8B A structural block diagram of an image forming apparatus and a server provided in an embodiment of this application; Figure 9 A schematic flowchart illustrating another image forming method provided in an embodiment of this application; Figure 10 A schematic flowchart illustrating another image forming method provided in an embodiment of this application; Figure 11 A schematic flowchart illustrating another image forming method provided in an embodiment of this application; Figure 12 A schematic flowchart illustrating another image forming method provided in an embodiment of this application; Figure 13 A schematic flowchart illustrating another image forming method provided in an embodiment of this application; Figure 14 A schematic flowchart illustrating another image forming method provided in an embodiment of this application; Figure 15 A schematic flowchart illustrating another image forming method provided in an embodiment of this application; Figure 16 This is a schematic diagram of the structure of an image forming apparatus provided in an embodiment of this application; Figure 17 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation

[0033] To better understand the technical solution of this application, the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0034] It should be understood that the described embodiments are merely some, not all, of the embodiments in this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0035] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. The singular forms “a,” “the,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0036] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0037] See Figure 1 This is a schematic diagram illustrating an application scenario provided by an embodiment of this application. For example... Figure 1 As shown, the application scenario includes: a terminal device 101, an image forming apparatus 102, and a server 103. The terminal device 101 and the image forming apparatus 102 are interconnected via a communication network for information transmission; the image forming apparatus 102 and the server 103 are also interconnected via a communication network for information transmission.

[0038] Among them, terminal devices 101 include, but are not limited to, desktop computers, laptop computers, networked computers, handheld computers, personal digital assistants (PDAs), internet-enabled mobile phones, smartphones, pagers, digital capture devices (e.g., digital cameras and camcorders), internet devices, e-books, information boards, and digital or network boards.

[0039] The image forming apparatus 102 has functions including, but not limited to, printing, scanning, copying, and faxing. For example, the image forming apparatus 102 can be one of the following product types: a single-function printer, an image forming apparatus with only printing capabilities; a multifunction printer, an image forming apparatus with printing, copying, scanning, and / or faxing capabilities, and the ability to selectively set the number of paper trays; a digital multifunction printer, based on copying capabilities, with standard or optional printing, scanning, and faxing functions, using digital principles and laser printing for document output, allowing for image and text editing as needed, possessing a large-capacity paper tray, high memory, large hard drive, strong network support, and multitasking capabilities.

[0040] The communication network between terminal device 101 and image forming apparatus 102, and / or between image forming apparatus 102 and server 103, can be a local area network (LAN) or a wide area network (WAN) relayed through a relay device. When the communication network is a LAN, for example, it can be a short-range communication network such as a Wi-Fi hotspot network, a Wi-Fi P2P network, a Bluetooth network, a Zigbee network, or a near-field communication (NFC) network. When the communication network is a WAN, for example, it can be a 3rd generation wireless telephone technology (3G) network, a 4th generation mobile communication technology (4G) network, a 5th generation mobile communication technology (5G) network, a 6th generation mobile communication technology (6G) network, a future public land mobile network (PLMN), or the Internet.

[0041] Artificial intelligence (AI) is a scientific technology that enables machines to simulate human intelligence. With the development of AI technology, some image forming devices can be used in conjunction with AI tools to meet users' personalized and diverse needs. For example, users may need to use AI tools to process the target image and the document to be processed acquired by the image forming device. This could include inserting the target image into the document or replacing the target image with the background of the document.

[0042] To address the aforementioned user needs, one implementation scheme in the relevant technology is as follows: First, the user sends the document to be processed to the image forming apparatus via a terminal device; then, the image forming apparatus uploads the obtained target image and the document to be processed together to the server; finally, the server uses artificial intelligence tools to process the target image and the document to be processed accordingly, obtains the processed document, and sends the processed document to the user's terminal device.

[0043] The above-mentioned implementation schemes in the related technologies have at least the following disadvantages: 1) They cannot intelligently control the process of the image forming apparatus acquiring the target image based on artificial intelligence tools, resulting in a low level of intelligence; 2) The target image and the document to be processed obtained by the image forming apparatus need to be uploaded to the server, resulting in poor security.

[0044] To address the aforementioned issues, this application provides a solution where an image acquisition instruction is generated using a large language model. This instruction is then used to execute the operation of acquiring the target image, achieving intelligent control of the target image acquisition process and improving its overall intelligence. Furthermore, processing the target image and the document to be processed is performed at the image forming apparatus. Since the image forming apparatus does not need to upload the acquired target image and the document to be processed to a server, security is improved. The specific implementation will be described in detail below.

[0045] See Figure 2 This is a schematic flowchart illustrating an image forming method provided in an embodiment of this application. This method can be applied to... Figure 1 The image forming apparatus shown, such as Figure 2 As shown, it mainly includes the following steps.

[0046] Step S201: Obtain the document to be processed and its description information, which includes the image acquisition strategy and the document processing strategy.

[0047] In this embodiment of the application, the document to be processed can be a Word document, a PDF document, a PPT document, etc., and this embodiment of the application does not impose specific restrictions on the format of the document to be processed.

[0048] In this embodiment, the descriptive information is natural language descriptive information that can be recognized by a large language model. Correspondingly, the image acquisition strategy is an image acquisition method described in natural language; the document processing strategy is a document processing method described in natural language. In specific implementations, the descriptive information can be text information or voice information, etc., and this embodiment does not impose specific limitations on the specific form of the descriptive information.

[0049] In one possible implementation, the image forming apparatus can acquire the document to be processed and the descriptive information through a terminal device. Specifically, acquiring the document to be processed and the descriptive information includes acquiring the document to be processed and the descriptive information sent by the terminal device.

[0050] For example, in Figure 3 In the illustrated application scenario, the user opens the image forming application on the terminal device and enters the image forming interface shown in 3A. As shown in 3A, the image forming interface of the terminal device includes an intelligent processing control 301. If the user wishes to perform intelligent processing, they can trigger the intelligent processing control 301 to enter the intelligent processing dialog interface shown in 3B. In the intelligent processing dialog interface, the terminal device and the image forming apparatus can engage in AI conversation.

[0051] As shown in 3B, the intelligent processing dialog interface includes an information input box 302, a document selection control 303, and a message sending control 304. If the user wishes to select a document to be processed from the documents stored in the terminal device, the document selection control 303 can be triggered, entering the document selection interface shown in 3C.

[0052] As shown in 3C, the document selection interface includes icons for multiple documents. If the user wishes to process document two, they can trigger the icon corresponding to document two, sending document two to the image forming device and entering the intelligent processing dialog interface shown in 3D. Document two can be understood as document 306 to be processed.

[0053] As shown in 3D, the intelligent processing dialog interface displays a document 306 (i.e., document two) to be sent to the terminal device. Furthermore, the user can enter corresponding description information 305 in the information input box 302, and after completing the input of the description information 305, trigger the message sending control 304 to enter the intelligent processing dialog interface shown in 3E.

[0054] As shown in 3E, the intelligent processing dialog interface displays the description information 305 sent to the terminal device. Thus, it can be understood that the user sends the document to be processed 306 and the description information 305 to the image forming apparatus via the terminal device.

[0055] It should be noted that the "trigger" involved in the embodiments of this application can be an operation such as a single click, double click, or long press, which will not be specifically described below.

[0056] In one possible implementation, the document to be processed and / or the descriptive information may also be data stored in the image forming apparatus, or data obtained by the image forming apparatus by other means (e.g., the image forming apparatus may obtain the document to be processed and / or the descriptive information by scanning). This application embodiment does not impose specific limitations on this.

[0057] Step S202: Based on the description information, determine the image acquisition instructions and document processing instructions generated by the large language model.

[0058] In this embodiment, the description information includes an image acquisition strategy and a document processing strategy. After the description information is input into a large language model, the large language model can output corresponding image acquisition instructions and document processing instructions. It can be understood that the image acquisition instructions are determined based on the image acquisition strategy in the description information; the document processing instructions are determined based on the document processing strategy in the description information.

[0059] In practice, the large language model can be deployed on a server, where it processes the descriptive information and generates corresponding image acquisition and document processing instructions; or it can be deployed in an image forming apparatus, where it processes the descriptive information and generates corresponding image acquisition and document processing instructions; or it can be deployed in both a server and an image forming apparatus, selectively using either the server or the image forming apparatus to process the descriptive information based on the actual application scenario or requirements.

[0060] For ease of explanation, in this embodiment, the large language model deployed on the server is referred to as the "first large language model," and the large language model deployed on the image forming apparatus is referred to as the "second large language model." The specific processes for processing descriptive information using the first large language model and the second large language model will be described in detail below.

[0061] Typically, due to the more abundant storage and computing resources of servers, large language models can be deployed on servers to improve prediction performance; that is, the first large language model can be a large language model. Conversely, due to the limited storage and computing resources of image forming apparatuses, smaller large language models can be deployed within the image forming apparatus to conserve these resources; that is, the second large language model can be a small language model. The size of the large language model can be quantified by parameters such as the number of model parameters, computational cost, storage and memory usage, and model structural dimensionality complexity.

[0062] It should be further noted that, depending on actual needs, those skilled in the art can also deploy small large language models in the server; and / or, deploy large large language models in the image forming apparatus. This application does not impose specific limitations on this.

[0063] Step S203: Execute the operation according to the image acquisition instruction to acquire the target image.

[0064] In this embodiment, the image acquisition instruction is an executable instruction of the image forming apparatus. Therefore, after the large language model generates the image acquisition instruction, the image forming apparatus can perform operations according to the image acquisition instruction to acquire the corresponding target image. Specifically, the image forming apparatus includes an image acquisition module, which can control the image acquisition module to perform corresponding image acquisition operations according to the image acquisition instruction, thereby obtaining the corresponding target image.

[0065] It is understandable that the image acquisition instructions generated by the large language model are used to acquire the target image. The large language model is used to analyze the specific image acquisition intention in the description information provided by the user, such as the requirements for the target image and the acquisition channel of the target image, so as to realize intelligent control of the target image acquisition process and improve the level of intelligence.

[0066] Step S204: Process the target image and the document to be processed according to the document processing instructions to obtain the processed document.

[0067] In this embodiment, the document processing instructions are executable instructions of the image forming apparatus. Therefore, after the large language model generates the document processing instructions, the image forming apparatus can perform corresponding processing on the target image and the document to be processed according to the document processing instructions to obtain the processed document. Specifically, the image forming apparatus includes a document processing module, which can control the document processing module to perform corresponding processing on the target image and the document to be processed according to the document processing instructions to obtain the processed document.

[0068] In one possible implementation, the target image and the document to be processed are processed according to document processing instructions to obtain a processed document. Specifically, this includes: inserting the target image into the document to be processed to obtain a processed document; or replacing the target image with the background of the document to be processed to obtain a processed document; or adjusting the document to be processed according to the style (font, font size, position information, scaling, etc.) of the target image to obtain a processed document; or extracting some or all information (e.g., text, symbols, patterns, etc.) from the target image and inserting the extracted information into the document to be processed to obtain a processed document.

[0069] It should be noted that the processing methods for the target image and the document to be processed described above are only some possible implementations listed in the embodiments of this application. In practical applications, users can also input other document processing strategies in the description information according to actual needs to perform other types of processing on the target image and the document to be processed. The embodiments of this application do not impose specific limitations on this.

[0070] In this embodiment, an image acquisition instruction is generated by a large language model, and the operation of acquiring the target image can be executed according to the image acquisition instruction, thereby realizing intelligent control of the target image acquisition process and improving the level of intelligence. Furthermore, processing of the target image and the document to be processed is performed at the image forming apparatus. Since the image forming apparatus does not need to upload the acquired target image and the document to be processed to a server, security can be improved.

[0071] In one possible implementation, the image acquisition instruction is a scanning instruction. Accordingly, performing an operation according to the image acquisition instruction to acquire the target image specifically includes: performing a scanning operation according to the scanning instruction to acquire the target image. That is, in this embodiment, the target image can be obtained through scanning. Specifically, the image acquisition module in the image forming apparatus includes a scanning module. The image forming apparatus can control the scanning module to perform corresponding scanning operations according to the scanning instruction to obtain a scanned image, and then use the obtained scanned image as the target image.

[0072] In this embodiment, scanning instructions are generated by a large language model, and corresponding scanning operations can be executed according to the scanning instructions, thereby achieving intelligent control of the scanning process and improving the level of intelligence. For example, precise control of scanning parameters (DPI, scanning area, brightness, contrast, etc.) can be achieved.

[0073] It should be further noted that, in addition to performing a scanning operation, the image acquisition instruction can also instruct the image forming apparatus to acquire the target image in other ways. For example, the image acquisition instruction may instruct the image forming apparatus to acquire an image that meets the requirements from images stored locally or on another device, and use it as the target image; or, the image acquisition instruction may instruct the image forming apparatus to retrieve an image that meets the requirements from a network, and use it as the target image, etc. The embodiments of this application do not impose specific limitations on this.

[0074] To facilitate understanding, the technical solutions provided in the embodiments of this application will be illustrated below with reference to some specific application scenarios.

[0075] For example, in Figure 4A In the application scenario shown, the user sends a document 401 to the image forming apparatus via a terminal device, which is "a two-page Word document." The description information 402 is "scan the paper in the document tray and insert the scanned image into the first page of the document." After receiving the description information 402, the image forming apparatus processes it using a large language model to generate corresponding scanning and document processing instructions. Based on the scanning instructions, the paper in the document tray is scanned to obtain the target image 403. Then, based on the document processing instructions, the target image 403 is inserted into the first page of the Word document, resulting in the processed document 404. Figure 4BAs shown.

[0076] For example, in Figure 5A In the application scenario shown, the user sends a document 501 to the image forming apparatus via a terminal device, which is described as "a 10-page PPT document." The description information 502 is "scan the paper in the document tray at 600 DPI with reduced contrast, and replace the background of all pages in the document with the scanned content." After receiving this description information 502, the image forming apparatus processes it using a large language model to generate corresponding scanning and document processing instructions. Based on the scanning instructions, the paper in the document tray is scanned at 600 DPI with reduced contrast to obtain the target image 503. Then, based on the document processing instructions, the background of all pages in the PPT document is replaced with the target image 503, resulting in the processed document 504. Figure 5B As shown.

[0077] For example, in Figure 6A In the application scenario shown, the user sends a document 601 to the image forming apparatus via a terminal device, which is "a one-page PDF document." The description information 602 is "scan the paper in the platen, extract the text from the scanned image, and insert it as the second paragraph of the document." After receiving the description information 602, the image forming apparatus processes it using a large language model to generate corresponding scanning and document processing instructions. Based on the scanning instructions, the paper in the platen is scanned to obtain a target image 603. Then, based on the document processing instructions, the text in the target image 603 is extracted and inserted as the second paragraph of the PDF document, resulting in the processed document 604. Figure 6B As shown.

[0078] For example, in Figure 6C In the application scenario shown, the user sends a document 605 to the image forming apparatus via a terminal device, which is described as "a one-page PDF document." The description information 606 is "scan the paper in the document tray and adjust the document according to the style of the scanned image." After receiving the description information 606, the image forming apparatus processes it using a large language model to generate corresponding scanning and document processing instructions. Based on the scanning instructions, the paper in the document tray is scanned to obtain the target image 607. Then, based on the document processing instructions, the text in the PDF document is adjusted according to the style (font, size, color, bold, underline, etc.) of the text in the target image 607, resulting in the processed document 608. Figure 6D As shown.

[0079] As described above, in the embodiments of this application, the description information can be processed by either a first large language model deployed in the server or a second large language model deployed in the image forming apparatus. These two implementation methods will be described separately below.

[0080] See Figure 7 This is a schematic flowchart illustrating another image forming method provided in an embodiment of this application. This method can be applied to... Figure 1 The application scenarios shown are as follows: Figure 7 As shown, it mainly includes the following steps.

[0081] Step S701: The image forming apparatus acquires the document to be processed and description information, the description information including the image acquisition strategy and the document processing strategy.

[0082] For details on this step, please refer to the description of step S201 above. For the sake of brevity, it will not be repeated here.

[0083] Step S702: The image forming apparatus sends description information to the server.

[0084] In this embodiment of the application, after the image forming apparatus obtains the description information, it sends the description information to the server so that the description information can be processed by the first large language model deployed on the server.

[0085] Step S703: The server uses the first major language model to process the description information and generate corresponding image acquisition instructions and document processing instructions.

[0086] In this embodiment of the application, after receiving the description information sent by the image forming apparatus, the server inputs the description information into the first large language model, and the first large language model can output corresponding image acquisition instructions and document processing instructions.

[0087] In one possible implementation, the description information also includes network retrieval description information, which instructs the server to perform a search on the network to obtain the target image. Therefore, when the server receives the network retrieval description information, it also performs a search on the network to obtain the target image.

[0088] In one possible implementation, after receiving the description information sent by the image forming apparatus, the server inputs the description information into a first language model. The first language model can output corresponding image acquisition instructions, document processing instructions, and web retrieval description instructions. The image acquisition instructions are determined based on the image acquisition strategy in the description information, the document processing instructions are determined based on the document processing strategy in the description information, and the web retrieval description instructions are determined based on the web retrieval description information.

[0089] In one possible implementation, the first language model can output only the corresponding image acquisition instructions and document processing instructions. In this case, the output after processing by the first language model is the image acquisition instructions, document processing instructions, and network retrieval description information in the description information.

[0090] Step S704: The server sends image acquisition instructions and document processing instructions to the image forming apparatus.

[0091] In this embodiment of the application, after generating image acquisition instructions and document processing instructions, the first language model sends the image acquisition instructions and document processing instructions to the image forming apparatus so that the image forming apparatus can perform the corresponding operations.

[0092] See Figure 8A This is a structural block diagram of an image forming apparatus and a server provided in an embodiment of this application. Figure 8A As shown, an MCP server is installed in the image forming apparatus, and an MCP client is installed in the server. The MCP client can send image acquisition instructions and document processing instructions generated by the first language model to the MCP server based on the MCP protocol.

[0093] It should be further noted that the MCP (Model Context Protocol) is a "universal interface" for the era of large models, allowing different large model applications to connect to and use various external tools, data sources, and APIs in a standard and secure manner. Of course, in addition to the MCP protocol, the image forming apparatus and the server can also exchange information based on other communication protocols, and this application embodiment does not impose specific limitations on this.

[0094] Furthermore, if the description information includes network search description information, after the server retrieves the target image from the network based on the network search description information, it also needs to send the retrieved target image to the image forming apparatus so that the image forming apparatus can perform corresponding processing.

[0095] Understandably, the server can directly retrieve the target image from the network based on the network retrieval description information in the description information, or it can retrieve the target image from the network based on the network retrieval description instructions processed by the first language model.

[0096] Step S705: The image forming apparatus performs an operation according to the image acquisition instruction to acquire the target image.

[0097] For example, in Figure 8AIn the implementation shown, after receiving an image acquisition command, the MCP server of the image forming apparatus calls the image acquisition model in the image forming apparatus to perform corresponding operations based on the MCP protocol to acquire the target image. It can be understood that when the image acquisition command is a scanning command, the MCP server can call the scanning module in the image forming apparatus to perform corresponding scanning operations based on the MCP protocol, and then use the obtained scanned image as the target image.

[0098] For further details regarding this step, please refer to the description of step S203 above. For the sake of brevity, these details will not be repeated here.

[0099] Step S706: The image forming apparatus processes the target image and the document to be processed according to the document processing instructions to obtain the processed document.

[0100] For example, in Figure 8A In the implementation shown, after receiving the document processing instruction, the MCP server of the image forming apparatus calls the document processing module in the image forming apparatus to process the target image and the document to be processed based on the MCP protocol, and obtains the processed document.

[0101] For further details regarding this step, please refer to the description of step S204 above. For the sake of brevity, these details will not be repeated here.

[0102] In this embodiment, a first large language model deployed on a server processes the descriptive information to generate image acquisition instructions and document processing instructions. Since the first large language model is typically a large language model, it requires a longer processing time but offers higher accuracy. Therefore, this solution is particularly suitable for application scenarios where timeliness is not critical but accuracy is paramount. Furthermore, image forming apparatuses typically can only communicate with servers when connected to a network; therefore, this solution is also particularly suitable for application scenarios where the image forming apparatus is connected to a network.

[0103] See Figure 9 This is a schematic flowchart illustrating another image forming method provided in an embodiment of this application. This method can be applied to... Figure 1 The application scenarios shown are as follows: Figure 9 As shown, it mainly includes the following steps.

[0104] Step S901: The image forming apparatus acquires the document to be processed and description information, the description information including image acquisition strategy, document processing strategy and network-attached retrieval description information.

[0105] In this embodiment of the application, in addition to the image acquisition strategy and document processing strategy, the description information also includes network-attached retrieval description information, which is used to instruct the server to retrieve files related to the target image in the network.

[0106] For further details regarding this step, please refer to the description of step S701 above. For the sake of brevity, these details will not be repeated here.

[0107] Step S902: The image forming apparatus sends description information to the server.

[0108] For details on this step, please refer to the description of step S702 above. For the sake of brevity, it will not be repeated here.

[0109] Step S903: The server uses the first major language model to process the description information and generate corresponding image acquisition instructions and document processing instructions.

[0110] In one possible implementation, after receiving the description information sent by the image forming apparatus, the server inputs the description information into a first language model. The first language model can output corresponding image acquisition instructions, document processing instructions, and network-attached retrieval description instructions. Among them, the image acquisition instructions are determined according to the image acquisition strategy in the description information, the document processing instructions are determined according to the document processing strategy in the description information, and the network-attached retrieval description instructions are determined according to the network-attached retrieval description information.

[0111] In one possible implementation, the first language model can output only the corresponding image acquisition instructions and document processing instructions. In this case, the output after processing by the first language model is the image acquisition instructions, document processing instructions, and network-attached retrieval description information in the description information. Step S904: The server sends image acquisition instructions and document processing instructions to the image forming apparatus.

[0112] For details on this step, please refer to the description of step S704 above. For the sake of brevity, it will not be repeated here.

[0113] Step S905: The image forming apparatus performs an operation according to the image acquisition instruction to acquire the target image.

[0114] For details on this step, please refer to the description of step S705 above. For the sake of brevity, it will not be repeated here.

[0115] Step S906: The image forming apparatus sends the target image to the server.

[0116] In this embodiment of the application, since the server needs to retrieve files related to the target image in the network, the image forming apparatus also needs to send the target image to the server after obtaining the target image.

[0117] Step S907: The server retrieves files related to the target image in the network based on the network retrieval description information and obtains the target file.

[0118] Specifically, after receiving the target image sent by the image forming apparatus, the server retrieves files related to the target image in the network based on the target image and network-attached retrieval description information, and obtains the target file.

[0119] For example, if the target image is item A, and the web search description is to search for 10 images related to the target image on the web, then the server will search for 10 images related to item A on the web.

[0120] It is understood that the server can directly search the network based on the network appended retrieval description information and the received target image in the description information, or it can search the network based on the network appended retrieval description instruction processed by the first language model and the received target image; the target file here can be an image related to the target image or a document related to the target image. This application embodiment does not impose specific restrictions on the format of the target file.

[0121] Step S908: The server sends the target file to the image forming apparatus.

[0122] In this embodiment of the application, after the server completes the retrieval of the target file, it sends the target file to the image forming apparatus so that the image forming apparatus can perform corresponding processing on the target file.

[0123] Step S909: The image forming apparatus processes the target image, the target file, and the document to be processed according to the document processing instructions to obtain the processed document.

[0124] In this embodiment of the application, the document to be processed can also be processed based on files retrieved from the network, thereby improving the intelligence of document processing, meeting the diverse document processing needs of users, and improving the user experience.

[0125] See Figure 10 This is a schematic flowchart illustrating another image forming method provided in an embodiment of this application. This method can be applied to... Figure 1 The application scenarios shown are as follows: Figure 10 As shown, it mainly includes the following steps.

[0126] Step S1001: The image forming apparatus acquires the document to be processed and description information, the description information including the image acquisition strategy and the document processing strategy.

[0127] For details on this step, please refer to the description of step S701 above. For the sake of brevity, it will not be repeated here.

[0128] Step S1002: The image forming apparatus sends description information and enhancement information to the server.

[0129] In one possible implementation, the enhancement information can be a description of the default image acquisition strategy of the image forming apparatus. For example, when the image acquisition strategy in the description information is to scan paper in the platen, the enhancement information can be the default scanning parameters of the image forming apparatus. It is understood that combining the default scanning parameters of the image forming apparatus to predict scanning instructions can improve the accuracy of the prediction results.

[0130] In one possible implementation, the image forming apparatus also includes a knowledge base, which can store various types of knowledge related to image forming. The enhancement information can be knowledge related to the descriptive information stored in the knowledge base of the image forming apparatus.

[0131] For example, in Figure 8B In the implementation shown, the image forming apparatus includes a knowledge base. After acquiring description information, the image forming apparatus can retrieve knowledge related to the description information from the knowledge base and send it to the server as enhancement information. For example, when the image acquisition strategy in the description information is to scan the paper in the platen, the predicted scanning parameters from the knowledge base can be used as enhancement information.

[0132] In addition to retrieving enhanced information describing the data from the knowledge base, the document processing module and image acquisition module in the image forming apparatus can also interact with the knowledge base to improve document processing and target image acquisition capabilities. For example, the document processing module can retrieve relevant knowledge from the knowledge base to improve document processing. Furthermore, the document processing module can send the processed document to the knowledge base for dynamic updates. Similarly, the image acquisition module can retrieve relevant knowledge from the knowledge base to acquire the target image. Furthermore, the image acquisition module can send the acquired target image to the knowledge base for updates.

[0133] It should be noted that the enhanced information may also include the image acquisition strategy description information of the image forming apparatus by default and the knowledge related to the description information in the knowledge base of the image forming apparatus. This will not be elaborated further in the embodiments of this application.

[0134] Step S1003: The server uses the first major language model to process the description information and augmentation information, and generates corresponding image acquisition instructions and document processing instructions.

[0135] In this embodiment, after receiving the description information and enhancement information, the server can input the description information and enhancement information into the first language model, and the first language model outputs the corresponding image acquisition instructions and document processing instructions.

[0136] For further details regarding this step, please refer to the description of step S703 above. For the sake of brevity, these details will not be repeated here.

[0137] Step S1004: The server sends image acquisition instructions and document processing instructions to the image forming apparatus.

[0138] Step S1005: The image forming apparatus performs an operation according to the image acquisition instruction to acquire the target image.

[0139] Step S1006: The image forming apparatus processes the target image and the document to be processed according to the document processing instructions to obtain the processed document.

[0140] For details regarding steps S1004-S1006, please refer to the description of steps S704-S706 above. For the sake of brevity, these details will not be repeated here.

[0141] In this embodiment, by using enhanced information from the knowledge base as a supplement to the descriptive information, the first language model can more accurately predict image acquisition instructions and document processing instructions, resulting in processed documents that better meet user expectations and improving user experience.

[0142] See Figure 11 This is a schematic flowchart illustrating another image forming method provided in an embodiment of this application. This method can be applied to... Figure 1 Image forming apparatus in the application scenarios shown, such as Figure 11 As shown, the method is in Figure 2 Based on the illustrated embodiment, step S202 specifically includes the following steps.

[0143] Step S2021: Use the second language model to process the description information and generate image acquisition instructions and document processing instructions.

[0144] In this embodiment, a second large language model deployed in the image forming apparatus processes the descriptive information to generate image acquisition instructions and document processing instructions. Since the second large language model is typically a smaller, more efficient large language model, it requires less time for data processing, but its accuracy is relatively lower than that of the first large language model. Therefore, this solution is particularly suitable for application scenarios that require a balance between timeliness and accuracy. Furthermore, because the second large language model is deployed in the image forming apparatus, its use is not affected even if the image forming apparatus is not connected to the network. Therefore, this solution is also particularly suitable for application scenarios where the image forming apparatus is not connected to the network.

[0145] In practical applications, the first or second language model can be selected to process the descriptive information based on the network connectivity of the image forming apparatus. Specifically, if the image forming apparatus is networked, then... Figure 7 The embodiment shown employs a first large language model deployed on a server to process the descriptive information; if the image forming apparatus is connected to the network, then... Figure 11 The solution in the illustrated embodiment uses a second large language model deployed in the image forming apparatus to process the descriptive information.

[0146] It should be noted that the specific details of the embodiments of this application can be found in the description above, and will not be repeated here for the sake of brevity.

[0147] In one possible implementation, a second major language model is used to process the descriptive information to generate image acquisition instructions and document processing instructions, including: processing the descriptive information and enhancement information using the second major language model to generate image acquisition instructions and document processing instructions; wherein, the enhancement information includes the image acquisition strategy description information of the image forming apparatus by default and / or knowledge related to the descriptive information in the knowledge base of the image forming apparatus.

[0148] For details on "enhanced information," please refer to the above introduction. For the sake of brevity, it will not be repeated here.

[0149] See Figure 12 This is a schematic flowchart illustrating another image forming method provided in an embodiment of this application. This method can be applied to... Figure 1 Image forming apparatus in the application scenarios shown, such as Figure 12 As shown, the method is in Figure 2 Based on the illustrated embodiment, the following steps are also included.

[0150] Step S1201: Send the processed document to the terminal device. The processed document is used for preview display on the terminal device's display interface.

[0151] In this embodiment of the application, after obtaining the processed document, the image forming apparatus sends the processed document to the user's terminal device for preview, so that the user can confirm the processed document and improve the user experience.

[0152] Furthermore, if the user finds that the processed document does not meet expectations after reviewing the preview, they can send new description information to the image forming apparatus so that the image forming apparatus can reprocess it, as will be explained in detail below.

[0153] See Figure 13 This is a schematic flowchart illustrating another image forming method provided in an embodiment of this application. This method can be applied to... Figure 1 Image forming apparatus in the application scenarios shown, such as Figure 13 As shown, the method is in Figure 12 Based on the illustrated embodiment, the following steps are also included.

[0154] Step S1301: Receive new description information sent by the terminal device, the new description information including description information of new image acquisition strategy and / or document processing strategy.

[0155] In this embodiment, if the previously generated processed document does not meet expectations, or if the user needs to regenerate the processed document for other reasons, the description information can be adjusted and then resent to the image forming apparatus. For ease of distinction, in this embodiment, the resent description information is referred to as "new description information".

[0156] It's important to explain that since the pending document was already sent during the previous generation of the processed document, and this document typically doesn't change, it's usually unnecessary to resend it when a new processed document needs to be generated. This avoids repetitive user actions and improves the user experience; it also saves on data processing volume.

[0157] Step S1302: Based on the new description information and historical description information, determine the new image acquisition instructions and new document processing instructions generated by the large language model.

[0158] For ease of distinction, in this embodiment of the application, the regenerated image acquisition instruction and document processing instruction are referred to as "new image acquisition instruction" and "new document processing instruction," respectively.

[0159] In this application's embodiments, "historical description information" refers to all or part of the description information previously received, excluding the new description information currently received by the image forming apparatus. In this application's embodiments, the large language model references historical description information for prediction, which can improve the accuracy of the prediction results.

[0160] Step S1303: Execute the operation according to the new image acquisition instruction to obtain a new target image.

[0161] For ease of distinction, in this embodiment of the application, the re-acquired target image is referred to as the "new target image".

[0162] Step S1304: Process the document to be processed according to the new document processing instructions and the new target image to obtain a new processed document.

[0163] For ease of distinction, in this embodiment of the application, the re-obtained processed document is referred to as the "new processed document".

[0164] Understandably, if the new processed document still does not meet the user's expectations, or if the user needs to regenerate the processed document again for other reasons, steps S1301-S1304 can be executed again.

[0165] In one possible implementation, after processing the document to be processed according to the new document processing instructions and the new target image to obtain a new processed document, the method further includes: updating enhancement information according to the new image acquisition instructions and the new document processing instructions; wherein the enhancement information includes the default image acquisition strategy description information of the image forming apparatus and / or knowledge related to the description information in the knowledge base of the image forming apparatus.

[0166] For details on "enhanced information," please refer to the above introduction. For the sake of brevity, it will not be repeated here.

[0167] Updating the enhanced information based on the new image acquisition instructions and the new document processing instructions can make the next processing of documents with similar instructions more in line with the user's needs, reducing the time spent regenerating documents that do not meet the user's expectations.

[0168] It should be noted that the specific details of steps S1301-S1304 can be found in the description above, and will not be repeated here for the sake of brevity.

[0169] In one possible implementation, the image forming apparatus can also receive a query instruction sent by the server, and after receiving the query instruction, send the functions supported by the image forming apparatus to the server.

[0170] For example, in Figure 8A and Figure 8BIn the implementation shown, the server can send a query command to the MCP server in the image forming apparatus via the MCP client. The MCP server can obtain the functions supported by the image forming apparatus based on the MCP protocol and feed back the supported functions to the MCP client, so that the server can know in advance the functions supported by the image forming apparatus.

[0171] Understandably, if the server is unaware of the functions supported by the image forming apparatus, the generated image acquisition and document processing instructions may contain some functions that the image forming apparatus does not support. In this case, even if the image acquisition and document processing instructions are sent to the image forming apparatus, the image forming apparatus will be unable to execute them. Therefore, by querying the functions supported by the image forming apparatus, the server can generate image acquisition and document processing instructions that are more compatible with the image forming apparatus.

[0172] In the above-described case, this application embodiment also provides a collaborative method for multiple image forming apparatuses, wherein multiple image forming apparatuses respectively receive a query instruction sent by a server, and after receiving the query instruction, each image forming apparatus sends its supported functions to the server, and the server selects the image forming apparatus that matches the user input description information for output.

[0173] See Figure 14 This is a schematic flowchart illustrating another image forming method provided in an embodiment of this application. This method can be applied to... Figure 1 The server in the application scenario shown, such as Figure 14 As shown, it mainly includes the following steps.

[0174] Step S1401: Receive description information sent by the image forming apparatus, the description information including image acquisition strategy and document processing strategy; Step S1402: Use the first language model to process the description information and generate corresponding image acquisition instructions and document processing instructions; Step S1403: Send an image acquisition instruction and a document processing instruction to the image forming apparatus; The image acquisition instruction is used to instruct the image forming apparatus to perform an operation corresponding to the image acquisition strategy; the document processing instruction is used to instruct the image forming apparatus to process the target image and the document to be processed according to the document processing strategy, so as to obtain the processed document.

[0175] In this embodiment, a first large language model deployed on a server processes the descriptive information to generate image acquisition instructions and document processing instructions. Since the first large language model is typically a large language model, it requires a longer processing time but offers higher accuracy. Therefore, this solution is particularly suitable for application scenarios where timeliness is not critical but accuracy is paramount. Furthermore, image forming apparatuses typically can only communicate with servers when connected to a network; therefore, this solution is also particularly suitable for application scenarios where the image forming apparatus is connected to a network.

[0176] For details regarding the embodiments of this application, please refer to the above text. Figure 7 The description of the embodiments shown is omitted here for the sake of brevity.

[0177] In one possible implementation, the description information further includes network retrieval description information, and the method further includes: searching the network based on the network retrieval description information to obtain the target image; sending the target image to the image forming apparatus, and a document processing instruction for instructing the image forming apparatus to process the target image and the document to be processed according to a document processing strategy.

[0178] For details regarding the embodiments of this application, please refer to the above text. Figure 7 The description of the embodiments shown is omitted here for the sake of brevity.

[0179] In one possible implementation, the network retrieval description information further includes network supplementary retrieval description information, and the method further includes: receiving a target image sent by the image forming apparatus; retrieving files related to the target image in the network according to the network supplementary retrieval description information, and obtaining the target file.

[0180] In this embodiment of the application, the document to be processed can also be processed based on files retrieved from the network, thereby improving the intelligence of document processing, meeting the diverse document processing needs of users, and improving the user experience.

[0181] For details regarding the embodiments of this application, please refer to the above text. Figure 9 The description of the embodiments shown is omitted here for the sake of brevity.

[0182] In one possible implementation, receiving description information sent by the image forming apparatus specifically includes: receiving description information and enhancement information sent by the image forming apparatus, wherein the enhancement information includes image acquisition strategy description information of the image forming apparatus by default and / or knowledge related to the description information in the knowledge base of the image forming apparatus; and processing the description information using a first major language model to generate corresponding image acquisition instructions and document processing instructions, specifically including: processing the description information and enhancement information using a first major language model to generate corresponding image acquisition instructions and document processing instructions.

[0183] In this embodiment, by using enhanced information from the knowledge base as a supplement to the descriptive information, the first language model can more accurately predict image acquisition instructions and document processing instructions, resulting in processed documents that better meet user expectations and improving user experience.

[0184] For details regarding the embodiments of this application, please refer to the above text. Figure 10 The description of the embodiments shown is omitted here for the sake of brevity.

[0185] In one possible implementation, sending an image acquisition instruction and a document processing instruction to the image forming apparatus includes: sending a query instruction to the image forming apparatus for querying functions supported by the image forming apparatus; receiving functions supported by the image forming apparatus sent by the image forming apparatus; and if the functions supported by the image forming apparatus match the scanning instruction and the document processing instruction, then sending the image acquisition instruction and the document processing instruction to the image forming apparatus.

[0186] Understandably, if the server is unaware of the functions supported by the image forming apparatus, the generated image acquisition and document processing instructions may contain some functions that the image forming apparatus does not support. In this case, even if the image acquisition and document processing instructions are sent to the image forming apparatus, the image forming apparatus will be unable to execute them. Therefore, by querying the functions supported by the image forming apparatus, the server can generate image acquisition and document processing instructions that are more compatible with the image forming apparatus.

[0187] In the above-described case, this application embodiment also provides a collaborative method for multiple image forming apparatuses, wherein the server can send query instructions to multiple image forming apparatuses respectively, and receive the supported functions sent by each image forming apparatus, and the server selects the image forming apparatus that matches the user input description information for output.

[0188] For details regarding the embodiments of this application, please refer to the description above. For the sake of brevity, further details will not be repeated here.

[0189] Furthermore, if the user finds that the processed document does not meet expectations after reviewing the preview, it is necessary to regenerate the image acquisition instructions and document processing instructions, which will be explained in detail below.

[0190] See Figure 15 This is a schematic flowchart illustrating another image forming method provided in an embodiment of this application. This method can be applied to... Figure 1 Image forming apparatus in the application scenarios shown, such as Figure 15 As shown, the method is in Figure 14 Based on the illustrated embodiment, the following steps are also included.

[0191] Step S1501: Receive new description information sent by the image forming apparatus, the new description information including description information of a new image acquisition strategy and / or document processing strategy; Step S1502: Use the first language model to process the new descriptive information and generate new image acquisition instructions and new document processing instructions; Step S1503: Send a new image acquisition instruction and a new document processing instruction to the image forming apparatus; The image acquisition instruction is used to instruct the image forming apparatus to perform corresponding scanning operations; the document processing instruction is used to instruct the image forming apparatus to perform corresponding processing on the document to be processed based on the target image.

[0192] It is understood that if the new processed document still does not meet the user's expectations, or if the user needs to regenerate the processed document for other reasons, steps S1501-S1504 can be executed again. For details regarding the embodiments of this application, please refer to the above description; for the sake of brevity, further elaboration will not be repeated here.

[0193] Corresponding to the above embodiments, this application also provides an image forming apparatus.

[0194] See Figure 16 This is a schematic diagram of the structure of an image forming apparatus provided in an embodiment of this application. Figure 16 As shown, the image forming apparatus 1600 includes a controller 1601, which is configured to perform some or all of the steps in the above method embodiments.

[0195] Corresponding to the above embodiments, this application also provides a server.

[0196] See Figure 17 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Figure 17As shown, the server 1700 may include a processor 1701, a memory 1702, and a communication unit 1703. These components communicate via one or more buses. Those skilled in the art will understand that the server structure shown in the figure does not constitute a limitation on the embodiments of this application. It may be a bus-type structure or a star-type structure, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0197] The communication unit 1703 is used to establish a communication channel, enabling the server to communicate with other devices.

[0198] Processor 1701 serves as the control center of the server, connecting various parts of the server via various interfaces and lines. It executes software programs and / or modules stored in memory 1702, and calls data stored in memory to perform various server functions and / or process data. The processor can be composed of integrated circuits (ICs), such as a single packaged IC or multiple packaged ICs with the same or different functions connected together. For example, processor 1701 may only include a central processing unit (CPU). In this embodiment, the CPU may have a single processing core or include multiple processing cores.

[0199] Memory 1702 is used to store the execution instructions of processor 1701. Memory 1702 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0200] When the execution instructions in memory 1702 are executed by processor 1701, the server 1700 is able to perform some or all of the steps in the above method embodiments.

[0201] Corresponding to the above embodiments, this application also provides a computer-readable storage medium, wherein the computer-readable storage medium may store a program, wherein when the program runs, it can control the device where the computer-readable storage medium is located to execute some or all of the steps in the above method embodiments. In specific implementation, the computer-readable storage medium may be a magnetic disk, an optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0202] Corresponding to the above embodiments, this application also provides a computer program product containing executable instructions that, when executed on a computer, cause the computer to perform some or all of the steps in the above method embodiments.

[0203] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, the simultaneous existence of A and B, or the existence of B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects have an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0204] Those skilled in the art will recognize that the units and algorithm steps described in the embodiments disclosed herein can be implemented using electronic hardware, computer software, or a combination of electronic hardware and software. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0205] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0206] In the several embodiments provided in this application, any function, if implemented as a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0207] The above description is merely a specific embodiment of this application. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application. The protection scope of this application should be determined by the protection scope of the claims.

Claims

1. An image forming method, characterized in that, Applied to an image forming apparatus, the method includes: Obtain the document to be processed and its description information, wherein the description information includes an image acquisition strategy and a document processing strategy; Based on the description information, image acquisition instructions and document processing instructions generated by the large language model are determined; wherein, the image acquisition instructions are determined according to the image acquisition strategy, and the document processing instructions are determined according to the document processing strategy; Execute the operation according to the image acquisition instruction to acquire the target image; The target image and the document to be processed are processed according to the document processing instructions to obtain the processed document.

2. The method according to claim 1, characterized in that, The image acquisition instruction is a scanning instruction, and the step of performing an operation according to the image acquisition instruction to acquire the target image includes: performing a scanning operation according to the scanning instruction to acquire the target image.

3. The method according to claim 1, characterized in that, The step of determining the image acquisition instructions and document processing instructions generated by the large language model based on the description information includes: The description information is sent to the server, which processes the description information using a first language model to generate corresponding image acquisition instructions and document processing instructions. Receive the image acquisition instruction and the document processing instruction sent by the server.

4. The method according to claim 3, characterized in that, The description information also includes network retrieval description information, and the server is further used to perform a search in the network based on the network retrieval description information to obtain the target image; The step of processing the target image and the document to be processed according to the document processing instructions to obtain the processed document further includes: Receive the target image sent by the server; process the target image and the document to be processed according to the document processing instructions to obtain the processed document.

5. The method according to claim 3, characterized in that, The description information also includes network-attached retrieval description information, which is used to indicate the retrieval of files related to the target image in the network based on the target image; After acquiring the target image by performing the operation according to the image acquisition instruction, the method further includes: The target image is sent to the server, and the server is further configured to retrieve files related to the target image in the network based on the network additional retrieval description information to obtain the target file; The step of processing the target image and the document to be processed according to the document processing instructions to obtain the processed document includes: Receive the target file sent by the server; The target image, the target file, and the document to be processed are processed according to the document processing instructions to obtain the processed document.

6. The method according to claim 3, characterized in that, The description information is sent to the server, and the server processes the description information using a first language model to generate corresponding image acquisition instructions and document processing instructions, including: The description information and enhancement information are sent to the server, which processes the description information and enhancement information using a first major language model to generate corresponding scanning instructions and document processing instructions. The enhanced information includes the default image acquisition strategy description information of the image forming apparatus and / or knowledge related to the description information in the knowledge base of the image forming apparatus.

7. The method according to claim 1, characterized in that, The step of determining the image acquisition instructions and document processing instructions generated by the large language model based on the description information includes: The description information is processed using a second major language model to generate the image acquisition instructions and the document processing instructions.

8. The method according to claim 7, characterized in that, The process of using a second major language model to process the descriptive information and generate the image acquisition instructions and the document processing instructions includes: The second major language model is used to process the descriptive information and enhanced information to generate image acquisition instructions and document processing instructions; The enhanced information includes the default image acquisition strategy description information of the image forming apparatus and / or knowledge related to the description information in the knowledge base of the image forming apparatus.

9. The method according to claim 1, characterized in that, The step of processing the document to be processed according to the document processing instructions and the target image in accordance with the document processing strategy to obtain the processed document includes: The target image is inserted into the document to be processed to obtain the processed document; Alternatively, the target image can be replaced with the background of the document to be processed to obtain the processed document; Alternatively, the document to be processed can be adjusted according to the style of the target image to obtain a processed document; Alternatively, some or all of the text can be extracted from the target image and inserted into the document to be processed to obtain the processed document.

10. The method according to claim 1, characterized in that, After obtaining the processed document, the following is also included: The processed document is sent to the terminal device, and the processed document is used for preview display on the display interface of the terminal device.

11. The method according to claim 10, characterized in that, Also includes: Receive new description information sent by the terminal device, the new description information including description information of new image acquisition strategy and / or document processing strategy; Based on the new description information and the historical description information, new image acquisition instructions and new document processing instructions generated by the large language model are determined; The operation is executed according to the new image acquisition instruction to obtain a new target image; The document to be processed is processed according to the new document processing instructions and the new target image to obtain a new processed document.

12. The method according to claim 11, characterized in that, After processing the document to be processed according to the new document processing instructions and the new target image to obtain a new processed document, the method further includes: Update the enhanced information according to the new image acquisition instructions and the new document processing instructions; The enhanced information includes the default image acquisition strategy description information of the image forming apparatus and / or knowledge related to the description information in the knowledge base of the image forming apparatus.

13. The method according to claim 3, characterized in that, Also includes: Receive a query instruction sent by the server, the query instruction being used to query the functions supported by the image forming apparatus; Send the functions supported by the image forming apparatus to the server.

14. An image forming method, characterized in that, Applied to a server, the method includes: Receive description information sent by the image forming apparatus, the description information including an image acquisition strategy and a document processing strategy; The first major language model is used to process the descriptive information to generate corresponding image acquisition instructions and document processing instructions; Send the image acquisition instruction and the document processing instruction to the image forming apparatus; The image acquisition instruction is used to instruct the image forming apparatus to perform an operation corresponding to the image acquisition strategy; the document processing instruction is used to instruct the image forming apparatus to process the target image and the document to be processed according to the document processing strategy, thereby obtaining a processed document.

15. The method according to claim 14, characterized in that, The description information also includes network retrieval description information, and the method further includes: The target image is obtained by searching the network based on the network retrieval description information. The target image is sent to the image forming apparatus, and the document processing instruction is used to instruct the image forming apparatus to process the target image and the document to be processed according to the document processing strategy.

16. The method according to claim 14, characterized in that, The description information also includes network-attached retrieval description information, and the method further includes: Receive the target image sent by the image forming apparatus; Based on the network additional retrieval description information, files related to the target image are retrieved in the network to obtain the target file.

17. The method according to claim 14, characterized in that, The receiving of description information sent by the image forming apparatus includes: receiving description information and enhancement information sent by the image forming apparatus, wherein the enhancement information includes the image acquisition strategy description information of the image forming apparatus by default and / or knowledge related to the description information in the knowledge base of the image forming apparatus; The step of processing the descriptive information using the first major language model to generate corresponding image acquisition instructions and document processing instructions includes: processing the descriptive information and the enhanced information using the first major language model to generate corresponding image acquisition instructions and document processing instructions.

18. The method according to claim 14, characterized in that, Sending the image acquisition instruction and the document processing instruction to the image forming apparatus includes: Send a query command to the image forming apparatus, the query command being used to query the functions supported by the image forming apparatus; Receive the functions supported by the image forming apparatus sent by the image forming apparatus; If the functions supported by the image forming apparatus match the image acquisition instruction and the document processing instruction, then the image acquisition instruction and the document processing instruction are sent to the image forming apparatus.

19. The method according to claim 14, characterized in that, Also includes: Receive new description information sent by the image forming apparatus, the new description information including description information of new image acquisition strategy and / or document processing strategy; The new descriptive information is processed using the first major language model to generate new image acquisition instructions and new document processing instructions; Send the new image acquisition instruction and the new document processing instruction to the image forming apparatus; The image acquisition instruction is used to instruct the image forming apparatus to perform a corresponding scanning operation; the document processing instruction is used to instruct the image forming apparatus to perform corresponding processing on the document to be processed based on the target image.

20. An image forming apparatus, characterized in that, include: A controller configured to perform the method of any one of claims 1-13.

21. A server, characterized in that, include: processor; Memory; And a computer program, wherein the computer program is stored in the memory, and when the computer program is executed by the processor, causes the server to perform the method of any one of claims 14-19.

22. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1-19.

23. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1-19.