Method and device for generating operation manual, computer device and storage medium

CN116090415BActive Publication Date: 2026-09-08INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310177428.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-28
Publication Date
2026-09-08
Estimated Expiration
2043-02-28

AI Technical Summary

Technical Problem

通常工作人员需熟悉应用程序的操作过程,并手动记录下每个操作步骤,形成操作手册,整个过程将耗费较多时间,存在效率较低的问题

Benefits of technology

[0058] The aforementioned method, apparatus, computer device, storage medium, and computer program product for generating operation manuals generate operation images corresponding to trigger operations during user operation of the application, and then use a translation strategy matching the operation type to translate the operation images to obtain translated content. Based on the translated content, operation manuals are generated automatically. This achieves the goal of automatically generating operation manuals when users operate the application, avoiding the process of users manually recording operation steps, thus saving operation manual writing time and improving operation manual writing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116090415B_ABST
    Figure CN116090415B_ABST
Patent Text Reader

Abstract

The application relates to a method and device for generating an operation manual, computer equipment, a storage medium and a computer program product, and relates to the technical field of artificial intelligence, and can be applied to the field of financial technology or other fields. The method comprises the following steps: monitoring a triggering operation of a user in a display interface of a target application program, and generating an operation image corresponding to the triggering operation based on current interface information of the display interface; performing classification processing on the operation image by using an image classification algorithm to obtain an operation type corresponding to the operation image; performing translation processing on the operation image by using a translation strategy matched with the operation type to obtain target translation content corresponding to the operation image; and generating an operation manual of the target application program based on the target translation content. The method can improve the generation efficiency of the operation manual.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, computer device, storage medium, and computer program product for generating an operation manual. Background Technology

[0002] The application development lifecycle generally includes five basic stages: design, development, testing, acceptance, and promotion. With the increasing popularity of mobile applications, the operation of mobile applications on the market has become increasingly complex. Therefore, when promoting applications, it is often necessary to provide user manuals to guide users in correctly using the various functions of the application.

[0003] In related technologies, application user manuals are typically written manually by staff (such as marketing personnel). This usually requires staff to be familiar with the application's operation and manually record each step to create the manual, a process that is time-consuming and inefficient. Summary of the Invention

[0004] Therefore, it is necessary to provide a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for generating operation manuals that can improve the efficiency of operation manual writing, in order to address the above-mentioned technical problems.

[0005] Firstly, this application provides a method for generating an operation manual. The method includes:

[0006] Monitor user actions triggered on the target application's display interface, and generate an action image corresponding to the triggered action based on the current interface information of the display interface;

[0007] The operation image is classified using an image classification algorithm to obtain the operation type corresponding to the operation image;

[0008] A translation strategy matching the operation type is used to translate the operation image to obtain the target translated content corresponding to the operation image;

[0009] Based on the target translated content, an operation manual for the target application is generated.

[0010] In one embodiment, generating the operation image corresponding to the trigger operation based on the current interface information of the display interface includes:

[0011] From the current interface information of the display interface, determine the trigger information corresponding to the trigger operation; the trigger information includes at least one of trigger object location information and trigger operation location information.

[0012] Based on the trigger information, an image marker is added to the current interface information of the display interface to obtain the operation image corresponding to the trigger operation.

[0013] In one embodiment, after generating the operation image corresponding to the trigger operation based on the current interface information of the display interface, the method further includes:

[0014] Based on the triggering order of the triggering operations, generate image index information of the operation image corresponding to the triggering operation;

[0015] The step of generating the operation manual for the target application based on the target translated content includes:

[0016] When there are multiple operation images, the target translation content corresponding to each operation image is sorted according to the image index information of each operation image, and an operation manual for the target application is generated based on the sorted target translation content.

[0017] In one embodiment, the step of employing a translation strategy matching the operation type to translate the operation image and obtain the target translated content corresponding to the operation image includes:

[0018] Based on a content extraction strategy that matches the operation type, content information corresponding to the triggering operation is extracted from the operation image; the content information includes content information of the triggering object corresponding to the triggering operation.

[0019] The target translation content corresponding to the operation image is generated based on the translation content template information matching the operation type and the content information corresponding to the trigger operation.

[0020] In one embodiment, extracting the content information corresponding to the triggering operation from the operation image based on a content extraction strategy matching the operation type includes:

[0021] When the operation type is an input operation type, the operation image is input to the image recognition model to obtain the target input box name image and the target input box image corresponding to the trigger operation in the operation image;

[0022] The target input box name image and the target input box image are subjected to text recognition to obtain the input box name and input content information.

[0023] In one embodiment, extracting the content information corresponding to the triggering operation from the operation image based on a content extraction strategy matching the operation type includes:

[0024] When the operation type is a text click operation, the clicked text information corresponding to the trigger operation in the operation image is identified;

[0025] When the operation type is an icon click operation, extract the click icon image corresponding to the trigger operation from the operation image.

[0026] In one embodiment, the method further includes:

[0027] Displays the chapter information entry interface;

[0028] In response to the user's completion of input on the chapter information input interface, a chapter image containing chapter information is obtained;

[0029] Identify the chapter information contained in the chapter image to obtain the target translated content corresponding to the chapter image;

[0030] The step of generating the operation manual for the target application based on the target translated content includes:

[0031] Based on the target translation content corresponding to the operation image and the target translation content corresponding to the chapter image, an operation manual for the target application is generated.

[0032] Secondly, this application also provides an apparatus for generating an operation manual. The apparatus includes:

[0033] The monitoring module is used to monitor the user's triggering operations on the display interface of the target application, and generate an operation image corresponding to the triggering operation based on the current interface information of the display interface.

[0034] The classification module is used to classify the operation image using an image classification algorithm to obtain the operation type corresponding to the operation image;

[0035] The translation module is used to translate the operation image using a translation strategy that matches the operation type, so as to obtain the target translated content corresponding to the operation image;

[0036] The first generation module is used to generate the operation manual of the target application based on the target translated content.

[0037] In one embodiment, the monitoring module is specifically used for:

[0038] In the current interface information of the display interface, determine the trigger information corresponding to the trigger operation; the trigger information includes at least one of trigger object location information and trigger operation location information; based on the trigger information, add image markers to the current interface information of the display interface to obtain the operation image corresponding to the trigger operation.

[0039] In one embodiment, the device further includes:

[0040] The second generation module is used to generate image index information of the operation image corresponding to the triggering operation based on the triggering order of the triggering operation;

[0041] The first generation module is specifically used for:

[0042] When there are multiple operation images, the target translation content corresponding to each operation image is sorted according to the image index information of each operation image, and an operation manual for the target application is generated based on the sorted target translation content.

[0043] In one embodiment, the translation module is specifically used for:

[0044] Based on a content extraction strategy that matches the operation type, content information corresponding to the triggering operation is extracted from the operation image; the content information includes trigger object content information corresponding to the triggering operation; target translated content corresponding to the operation image is generated based on translated content template information that matches the operation type and the content information corresponding to the triggering operation.

[0045] In one embodiment, the translation module is specifically used for:

[0046] When the operation type is an input operation type, the operation image is input to the image recognition model to obtain the target input box name image and the target input box image corresponding to the trigger operation in the operation image; text recognition is performed on the target input box name image and the target input box image respectively to obtain the input box name and input content information.

[0047] In one embodiment, the translation module is specifically used for:

[0048] When the operation type is a text click operation, the clicked text information corresponding to the trigger operation in the operation image is identified; when the operation type is an icon click operation, the clicked icon image corresponding to the trigger operation in the operation image is extracted.

[0049] In one embodiment, the device further includes:

[0050] The display module is used to display the chapter information input interface;

[0051] The acquisition module is used to acquire a chapter image containing chapter information in response to the user's completion of the entry operation on the chapter information entry interface;

[0052] The recognition module is used to recognize the chapter information contained in the chapter image and obtain the target translated content corresponding to the chapter image;

[0053] The first generation module is specifically used for:

[0054] Based on the target translation content corresponding to the operation image and the target translation content corresponding to the chapter image, an operation manual for the target application is generated.

[0055] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the method described in the first aspect.

[0056] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the method described in the first aspect.

[0057] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.

[0058] The aforementioned method, apparatus, computer device, storage medium, and computer program product for generating operation manuals generate operation images corresponding to trigger operations during user operation of the application, and then use a translation strategy matching the operation type to translate the operation images to obtain translated content. Based on the translated content, operation manuals are generated automatically. This achieves the goal of automatically generating operation manuals when users operate the application, avoiding the process of users manually recording operation steps, thus saving operation manual writing time and improving operation manual writing efficiency.

[0059] Furthermore, this method determines the corresponding operation type by classifying the operation images, which can be achieved based solely on the information displayed on the front-end interface of the application. It does not require accessing the operation information in the underlying back-end data of the application, thus avoiding the occupation of the application's back-end resources. The operation manual generation process does not affect the application's performance, and it can also be applied to scenarios where the application's back-end data cannot be accessed, making its application scenarios more extensive. Attached Figure Description

[0060] Figure 1 This is a flowchart illustrating a method for generating an operation manual in one embodiment;

[0061] Figure 2 This is a flowchart illustrating the process of generating an operation image in one embodiment;

[0062] Figure 3 This is a flowchart illustrating the process of obtaining the target translated content in one embodiment;

[0063] Figure 4 This is a flowchart illustrating the method for generating an operation manual in another embodiment;

[0064] Figure 5 This is a structural block diagram of an apparatus for generating an operation manual in one embodiment;

[0065] Figure 6 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0066] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0067] First, before introducing the technical solutions of the embodiments of this application, the technical background or evolution of the embodiments of this application will be introduced. With the popularization of mobile applications, the operation of mobile applications on the market has become increasingly complex. Therefore, when promoting applications, it is often necessary to provide a detailed operation manual to guide users in correctly using the various functions of the application. In related technologies, the operation manuals of applications are generally manually written by staff (such as promotion personnel). Typically, staff need to be familiar with the operation process of the application and manually record each operation step. They usually record the textual description of the operation step (such as clicking a button / control) and the corresponding screenshot of the interface, and then compile the operation steps into an operation manual. The entire process consumes a lot of time and manpower, resulting in low efficiency.

[0068] Against this backdrop, the applicant, through long-term research and development and experimental verification, proposes a method for generating operation manuals. This method generates operation images corresponding to triggered operations during user (e.g., testers of acceptance software) operation of the application. Then, a translation strategy matching the operation type is used to translate the operation images, obtaining translated content. The operation manual is then generated based on this translated content. This allows for automatic generation of operation manuals as users operate the application, avoiding the need for users to manually record operation steps, thus saving operation manual writing time and improving efficiency. Furthermore, this method determines the corresponding operation type by classifying operation images, relying solely on information from the application's front-end display interface. It does not require access to operation information in the application's back-end underlying data, thus avoiding the consumption of back-end resources and not affecting application performance during operation manual generation. It is also applicable to scenarios where back-end data is inaccessible, broadening its application scope. Additionally, it should be noted that the applicant has devoted significant creative effort to discovering the technical problem and the technical solutions described in the following embodiments.

[0069] In one embodiment, such as Figure 1 As shown, a method for generating an operation manual is provided. This method can be applied to a terminal, and it is understood that it can also be applied to a server, and furthermore, to a system including both a terminal and a server, and is implemented through interaction between the terminal and the server. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. In this embodiment, the method includes the following steps:

[0070] Step 101: Monitor the user's triggered actions on the target application's display interface, and generate an operation image corresponding to the triggered actions based on the current interface information of the display interface.

[0071] The target application refers to the application from which the user manual is to be generated. The target application can be installed on the terminal, and its interface can be displayed.

[0072] In implementation, the terminal can monitor user actions on the target application's display interface. Upon detecting any user action, it can generate an operation image corresponding to that action based on the current interface information. This operation image is used for subsequent operation type classification and translation. It can include an image of the triggering object corresponding to the action in the current display interface; that is, it can be a partial image of the triggering object within the current interface, or it can be a complete image of the current interface. The selection range of the operation image can be set according to actual needs, such as setting different image capture methods for different types of triggering actions. The triggering object can be an icon, text button, input box, or display page in the interface. Different triggering actions correspond to different triggering objects. For example, a click action might correspond to an icon or text button, an input action might correspond to an input box, and a swipe action might correspond to a display page.

[0073] Step 102: Use an image classification algorithm to classify the operation image to obtain the operation type corresponding to the operation image.

[0074] In implementation, after the terminal receives the operation image corresponding to the triggered operation, it can use an image classification algorithm to classify the operation image to obtain the operation type corresponding to the operation image. Operation types can include click operations, input operations, swipe operations, etc. Optionally, click operations can be further divided into text click operations, icon click operations, etc., and the specific granularity of classification can be set as needed. Since operation images of different operation types have different feature information—for example, click operation images generally contain image information of clicked buttons or other controls, while input operation images generally contain image information of input boxes (rectangular boxes, rounded boxes, straight boxes, etc.)—image classification algorithms can be used to classify the operation images. For example, an image classification model can be pre-built based on an image classification algorithm, and the model can be trained using training sample images labeled with operation types. The trained image classification model can then be used to classify the operation images to obtain their corresponding operation types.

[0075] Step 103: Using a translation strategy that matches the operation type, the operation image is translated to obtain the target translated content corresponding to the operation image.

[0076] In implementation, after classifying the operation images, the terminal can employ a translation strategy matching the operation type of the image to translate it and obtain the target translated content. Different operation types correspond to different translation strategies, and a pre-established correspondence between operation types and translation strategies can be established. The translation strategy can identify key information in the operation image, such as content information and interface identifiers, and then obtain the target translated content based on this key information and template information. For example, for an operation image of a click operation type, the content information (text or icon) of the clicked object can be identified, and the corresponding target translated content could be "Click on 'Object A'".

[0077] Step 104: Generate the user manual for the target application based on the target translated content.

[0078] In implementation, the terminal can generate an operation manual by translating the target content corresponding to the operation image according to a preset document generation format. Understandably, if multiple operation images are involved, the translated content corresponding to each operation image can be sequentially concatenated according to the operation order and then used to generate an operation manual according to the document generation format. Furthermore, the operation manual can also include operation images; that is, the target translated content corresponding to the operation image can be written into a document according to a preset format, and operation images can be inserted into the document to obtain an operation manual containing textual descriptions and image information of the operation steps, making it easier for users to read and understand.

[0079] In the aforementioned method for generating user manuals, operation images corresponding to triggered operations are generated during user interaction with the application. These images are then translated using a translation strategy matching the operation type, resulting in translated content. The user manual is then generated based on this translated content. This method automatically generates user manuals as the user interacts with the application, eliminating the need for manual recording of steps and saving time and increasing efficiency in manual creation. Furthermore, this method determines the operation type by classifying the operation images, relying solely on information from the application's front-end display interface. It does not require access to the application's back-end data, thus avoiding the consumption of back-end resources and maintaining application performance during manual generation. It is also applicable to scenarios where back-end data is inaccessible, broadening its application scope.

[0080] In one embodiment, such as Figure 2 As shown, the process of generating the operation image in step 101 specifically includes the following steps:

[0081] Step 201: Determine the trigger information corresponding to the trigger operation from the current interface information displayed on the display interface.

[0082] The trigger information includes at least one of the trigger object location information and the trigger operation location information.

[0083] In implementation, when the terminal detects a user's trigger operation, it can take a screenshot to obtain the current interface image. Furthermore, the terminal can determine the location information of the trigger operation within the current interface (trigger operation location information), and then determine the trigger information based on this location information. For example, for click or input operations, the terminal can determine the corresponding trigger object (i.e., the graphical interface element in the current interface corresponding to the trigger operation's location) based on the trigger operation location information, and then use the location information of that trigger object as the trigger information. For swipe operations, the terminal can use the initial swipe location information and the swipe path information as the corresponding trigger information.

[0084] Step 202: Based on the trigger information, add an image marker to the current interface information of the display interface to obtain the operation image corresponding to the trigger operation.

[0085] In implementation, after the terminal determines the trigger information, it can add image markers to the current interface image based on the location information in the trigger information. For example, if the trigger information is the location information of the trigger object, the terminal can add image markers, such as adding marker boxes, to the image area in the current interface image that matches the location information of the trigger object, to mark the location of the trigger object (such as a clicked button or input box). If the trigger information is the location information of the trigger operation (such as the initial position information and swipe path information of a swipe operation), the initial swipe position and swipe path can be marked in the current interface image to obtain the operation image corresponding to the trigger operation.

[0086] In this embodiment, when a user's trigger operation is detected, a screenshot of the current interface is taken and the location of the trigger operation is marked to obtain an operation image. Based on the marking and screenshot content of the operation image, the operation image can be classified. This avoids obtaining operation information from the application's backend underlying data, which would reduce the application's performance. At the same time, it can ensure the accuracy of classification. Furthermore, the obtained operation image can also be used as the image display part of each operation step in the operation manual, which can more intuitively and clearly show the operation steps and facilitate user understanding.

[0087] In one embodiment, after generating the operation image in step 101, the method further includes generating image index information. The specific process is as follows: Based on the triggering order of the triggering operations, image index information of the operation images corresponding to the triggering operations is generated. Correspondingly, the process of generating the operation manual in step 104 specifically includes the following steps: When there are multiple operation images, the target translation content corresponding to each operation image is sorted according to the image index information of each operation image, and an operation manual for the target application is generated based on the sorted target translation content.

[0088] In implementation, users typically trigger multiple operations during use, resulting in multiple operation images. Based on the triggering order, image index information corresponding to each operation can be generated. This image index information can identify each operation image and also reflect their relative order. For example, the operation image obtained from the first triggering operation can be labeled "1," and the operation image obtained from the second triggering operation can be labeled "2." "1" and "2" are the image index information for the operation images. After identifying the target translation content corresponding to each operation image, the terminal can sort the target translation content according to the image index information to obtain target translation content ordered according to the triggering order. Then, the sorted target translation content can be written into a document line by line in order, generating an operation manual with accurate operation steps.

[0089] Optionally, the terminal can add information such as the image index of the operation image, the operation type, and the target translation content to the translation information table, and sort them according to the order reflected in the image index information. In one example, the translation information table is shown in Table 1. The terminal can read the target translation content in the translation information table one by one, and write the target translation content into a document according to the text format matching the operation type to generate an accurate operation manual. In addition, before generating the operation manual, the terminal can show the translation information table to the user for modification and confirmation, and then generate the operation manual based on the user-confirmed translation information table to further ensure the accuracy of the operation manual.

[0090] Table 1 Translation Information Table

[0091]

[0092]

[0093] Understandably, for swipe-based operation images, the terminal can identify the swipe operation type that reflects the swipe direction based on the markers added to the operation image (i.e., the initial swipe position and swipe path). Furthermore, it can categorize swipes into two main types based on whether the swipe position includes a progress bar: swipes with a progress bar and swipes without a progress bar. Based on the swipe direction and the presence or absence of a progress bar, swipe operation types can be further subdivided, such as left swipe - no progress bar, right swipe - no progress bar, up swipe - no progress bar, down swipe - no progress bar, left swipe - with progress bar, right swipe - with progress bar, up swipe - with progress bar, down swipe - with progress bar, etc. This allows the translation content template information corresponding to a specific swipe operation type to be used as the target translation content. For example, the translation content template information for a left swipe operation without a progress bar could be "swipe the screen to the left".

[0094] In this embodiment, image index information is generated for each operation image according to the triggering order, so that the target translation content corresponding to each operation image is arranged according to the image index information, resulting in an operation manual with accurate operation order. Furthermore, based on this method, while ensuring the accuracy of the order of each operation step in the operation manual, different terminals can execute the operation image generation step of step 101 and subsequent steps (image classification, translation, and manual generation steps 102 to 104) respectively. The two terminals only need to interact with the operation images, resulting in simple data interaction and high efficiency. For example, a first terminal (such as a mobile phone) can monitor the triggering operations of applications installed by the user on this terminal to obtain multiple operation images numbered according to the triggering order. Then, the first terminal can send each operation image to a second terminal (such as a computer), which performs subsequent classification and translation processing on each operation image and generates an accurate operation manual based on the index information of each operation image. Therefore, this method is applicable to situations where the first terminal is unable to run complex image classification algorithms or other programs due to performance limitations.

[0095] In one embodiment, such as Figure 3 As shown, the process of obtaining the target translated content in step 103 specifically includes the following steps:

[0096] Step 301: Based on the content extraction strategy that matches the operation type, extract the content information corresponding to the trigger operation from the operation image.

[0097] The content information includes the content information of the trigger object corresponding to the trigger operation. Different content extraction strategies can be pre-set for different operation types, allowing for the use of matching strategies to extract content information from images of different operation types. For example, for click operations, the text content of the clicked trigger object (such as a text button) can be extracted; for input operations, the input box identifier (such as a name) and the text content within the input box can be extracted.

[0098] Step 302: Generate the target translation content corresponding to the operation image based on the translation content template information that matches the operation type and the content information corresponding to the trigger operation.

[0099] In implementation, translation content template information corresponding to each operation type can be pre-set. The translation content template information can contain text reflecting the execution action that triggers the operation, so that users can refer to it after reading and execute the relevant actions to correctly trigger the corresponding operation. For example, the translation content template information for a click operation can be "Click 'A'", where "A" can be replaced with the content information obtained in step 301. If the triggering object corresponding to the click operation is a "Login" control, the content information can be extracted as "Login", and then "A" in the template information can be replaced with "Login", thus obtaining the target translation content as "Click 'Login'".

[0100] In this embodiment, by extracting the content information corresponding to the triggered operation from the operation image, the translated content corresponding to the operation image can be generated based on the translated content template information and the extracted content information. The translated content includes the content information of the operation action and the operation object, which makes it easier for users to read and understand the operation steps, thereby guiding users to use the application correctly.

[0101] In one embodiment, the process of extracting content information in step 301 above specifically includes the following steps: when the operation type is an input operation type, the operation image is input to the image recognition model to obtain the target input box name image and the target input box image corresponding to the trigger operation in the operation image; text recognition is performed on the target input box name image and the target input box image respectively to obtain the input box name and input content information.

[0102] In implementation, the operation type can include input operation type. For operation images of input operation type, the terminal can input the operation image into the image recognition model to further identify the target input box name image and the target input box image corresponding to the trigger operation in the operation image. In practical applications, input boxes can have various styles: square boxes, rounded square boxes (squares with rounded corners), straight boxes, etc. Input boxes are generally distributed adjacent to their names, and their distribution can vary, such as the input box name being to the left, above, right, or below the input box. Therefore, training sample images containing different styles of input boxes and different name distributions can be used to train the image recognition model to improve recognition accuracy and robustness. The trained image recognition model can be used to identify the input box image and input box name image in the operation image (or at the marked position in the operation image). Then, the terminal can use an OCR (Optical Character Recognition) model to perform text recognition on the input box name image and the input box image respectively, obtaining the input box name and input content information. These two types of extracted content information are used to replace the corresponding content in the translation content template information to obtain the target translation content. In one example, the translation content template information corresponding to the input operation type is "Enter 'B' in the 'A' input box", where "A" can be replaced with the extracted input box name and "B" can be replaced with the extracted input content information, thus obtaining the target translation content.

[0103] In this embodiment, for operation images of input operation types, the image recognition model identifies the input box image and input box name image corresponding to the trigger operation. Then, text recognition is performed on the two images respectively, which can accurately identify the input box name and input content corresponding to this trigger operation, and then generate translated content containing the input box name and input content so that the user can understand the operation steps more clearly.

[0104] In one embodiment, the operation type includes text click operation type and image click operation type. The process of extracting content information in step 301 specifically includes the following steps: when the operation type is text click operation type, identify the clicked text information corresponding to the operation in the operation image; when the operation type is icon click operation type, extract the clicked icon image corresponding to the operation in the operation image.

[0105] In implementation, text click operation type refers to operation type where the trigger object is a text control or text button control, while image click operation type refers to operation type where the trigger object is an icon control. For text click operation type images, the terminal can directly identify the text information contained in the trigger object (i.e., the clicked text information corresponding to the trigger operation) as the content information corresponding to that operation image. For icon click operation type images, the terminal can obtain the image of the trigger object (i.e., the image area marked on the operation image), which is the clicked icon image corresponding to the trigger operation, and serves as the content information corresponding to that operation image.

[0106] In this embodiment, the click operation can specifically include text click operation and icon click operation. For different click operation types, the text of the clicked object or the icon is used as content information to replace the corresponding content in the translated content template information and generate an operation manual. This results in an operation manual that matches the actual operation and is easy to read, which helps guide users to use the application correctly.

[0107] In one embodiment, such as Figure 4 As shown, the method also includes the acquisition and translation of chapter images, specifically including the following steps:

[0108] Step 401: Display the chapter information entry interface.

[0109] In implementation, the user manual can include chapter information and corresponding operation steps for each chapter to clearly demonstrate the application's usage process, facilitating user reading and understanding. To obtain chapter information, a chapter information entry plugin can be installed on the terminal, and this plugin control (such as a floating icon) can be displayed on the interface. Users can click on this plugin control to trigger a chapter information entry request, and the terminal can respond to this request by displaying the chapter information entry interface.

[0110] Step 402: In response to the user's completion of the entry operation on the chapter information entry interface, obtain the chapter entry page image containing the chapter information.

[0111] In implementation, users can enter chapter information on the chapter information entry interface. Chapter information may include the chapter number (e.g., Chapter 1), title, and a description of the chapter content. After entering the chapter information, users can trigger a completion action, such as clicking the "Complete" button on the chapter information entry interface. The terminal can respond to this completion action by capturing a screenshot of the current interface, obtaining an image of the chapter containing the chapter information. Understandably, after the user clicks the "Complete" button or the "Exit" button, the terminal can exit the chapter information entry interface and display the target application's interface, allowing the terminal to begin monitoring user actions triggered within the target application's display interface.

[0112] Step 403: Identify the chapter information contained in the chapter image to obtain the target translated content corresponding to the chapter image.

[0113] In implementation, the terminal can identify the chapter information contained in the chapter images, using it as the target translation content corresponding to the chapter entry page images. Therefore, based on the target translation content corresponding to the chapter images and the target translation content corresponding to the operation images, an operation manual containing chapter information and operation steps can be generated.

[0114] Understandably, the user manual can contain multiple chapters. Users can enter chapter information one by one, and after completing the entry for each chapter, trigger the operation steps under that chapter on the target application's display interface. Thus, the terminal can generate image index information for each chapter image and operation image according to the order of chapter information entry and the order of triggering operations. This information identifies each image and reflects the sequential order of the chapter and operation images. Then, based on the image index information of each chapter and operation image, the target translated content can be sorted to generate a user manual with accurate chapter order and accurate operation steps under each chapter. Furthermore, after obtaining the chapter and operation images, the terminal can use an image classification algorithm in step 102 to classify each image to determine the chapter and operation images, as well as the specific operation type of the operation image. Then, the terminal can use a translation strategy matching the image category to translate each image, obtaining the corresponding target translated content, thereby generating the user manual.

[0115] In this embodiment, the terminal can also display a chapter information input interface so that users can input chapter information. After the user completes the input, the terminal can obtain a chapter image containing the chapter information, and then identify the chapter information contained in the chapter image as the target translation content corresponding to the chapter image. Thus, based on the target translation content of the chapter image and the target translation content of the operation image, an operation manual containing chapter information and operation steps can be obtained, which is more convenient for users to read.

[0116] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0117] Based on the same inventive concept, this application also provides an apparatus for generating an operation manual to implement the above-described method for generating operation manuals. The solution provided by this apparatus is similar to the solution described in the above method; therefore, the specific limitations in one or more operation manual generation apparatus embodiments provided below can be found in the limitations of the operation manual generation method described above, and will not be repeated here.

[0118] In one embodiment, such as Figure 5 As shown, an operation manual generation device 500 is provided, including: a monitoring module 501, a classification module 502, a translation module 503, and a first generation module 504, wherein:

[0119] The monitoring module 501 is used to monitor the user's triggered operations on the display interface of the target application, and generate the operation image corresponding to the triggered operation based on the current interface information of the display interface.

[0120] The classification module 502 is used to classify the operation image using an image classification algorithm to obtain the operation type corresponding to the operation image.

[0121] The translation module 503 is used to translate the operation image using a translation strategy that matches the operation type, so as to obtain the target translated content corresponding to the operation image.

[0122] The first generation module 504 is used to generate an operation manual for the target application based on the target translated content.

[0123] In one embodiment, the monitoring module 501 is specifically used to: determine the trigger information corresponding to the trigger operation in the current interface information of the display interface; the trigger information includes at least one of trigger object location information and trigger operation location information; and add image markers to the current interface information of the display interface based on the trigger information to obtain the operation image corresponding to the trigger operation.

[0124] In one embodiment, the device further includes a second generation module, configured to generate image index information of the operation images corresponding to the triggering operations based on the triggering order of the triggering operations. The corresponding first generation module 504 is specifically configured to: when there are multiple operation images, sort the target translation content corresponding to each operation image according to the image index information of each operation image, and generate an operation manual for the target application based on the sorted target translation content.

[0125] In one embodiment, the translation module 503 is specifically used to: extract content information corresponding to the triggering operation from the operation image based on a content extraction strategy that matches the operation type; the content information includes the triggering object content information corresponding to the triggering operation; and generate target translated content corresponding to the operation image based on the translated content template information that matches the operation type and the content information corresponding to the triggering operation.

[0126] In one embodiment, the translation module 503 is specifically used to: when the operation type is an input operation type, input the operation image into the image recognition model to obtain the target input box name image and the target input box image corresponding to the trigger operation in the operation image; perform text recognition on the target input box name image and the target input box image respectively to obtain the input box name and input content information.

[0127] In one embodiment, the translation module 503 is specifically used to: when the operation type is a text click operation, identify the clicked text information corresponding to the trigger operation in the operation image; when the operation type is an icon click operation, extract the clicked icon image corresponding to the trigger operation in the operation image.

[0128] In one embodiment, the device further includes a display module, an acquisition module, and an identification module, wherein:

[0129] The display module is used to show the chapter information entry interface.

[0130] The acquisition module is used to acquire a chapter image containing chapter information in response to the user's completion of the entry operation on the chapter information entry interface.

[0131] The recognition module is used to identify the chapter information contained in the chapter image and obtain the target translated content corresponding to the chapter image.

[0132] Accordingly, the first generation module 504 is specifically used to: generate the operation manual of the target application based on the target translation content corresponding to the operation image and the target translation content corresponding to the chapter image.

[0133] Each module in the aforementioned manual generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0134] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 6 As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a method for generating an operation manual. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.

[0135] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0136] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0137] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0138] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0139] The method, apparatus, computer equipment, storage medium, and computer program products for generating operation manuals provided in this application relate to the field of artificial intelligence technology and can be used in the field of financial technology or other fields. This application does not limit the application field.

[0140] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0141] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0142] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0143] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for generating an operation manual, characterized in that, The method includes: Monitor user actions on the target application's display interface, determine trigger information corresponding to the action from the current interface information, the trigger information includes at least one of trigger object location information and trigger action location information, add image markers to the current interface information based on the trigger information to obtain the action image corresponding to the action; wherein, the target application refers to the application for which the operation manual is to be generated; Based on the triggering order of the triggering operations, generate image index information of the operation image corresponding to the triggering operation; The operation image is classified using an image classification algorithm to obtain the operation type corresponding to the operation image; A translation strategy matching the operation type is used to translate the operation image to obtain the target translated content corresponding to the operation image; Based on the target translated content, an operation manual for the target application is generated; wherein, when there are multiple operation images, the target translated content corresponding to each operation image is sorted according to the image index information of each operation image, and the operation manual for the target application is generated based on the sorted target translated content.

2. The method according to claim 1, characterized in that, The step of employing a translation strategy matching the operation type to translate the operation image and obtain the target translated content corresponding to the operation image includes: Based on a content extraction strategy that matches the operation type, content information corresponding to the triggering operation is extracted from the operation image; the content information includes content information of the triggering object corresponding to the triggering operation. The target translation content corresponding to the operation image is generated based on the translation content template information matching the operation type and the content information corresponding to the trigger operation.

3. The method according to claim 2, characterized in that, The content extraction strategy based on matching the operation type, which extracts the content information corresponding to the triggering operation from the operation image, includes: When the operation type is an input operation type, the operation image is input to the image recognition model to obtain the target input box name image and the target input box image corresponding to the trigger operation in the operation image; The target input box name image and the target input box image are subjected to text recognition to obtain the input box name and input content information.

4. The method according to claim 2, characterized in that, The content extraction strategy based on matching the operation type, which extracts the content information corresponding to the triggering operation from the operation image, includes: When the operation type is a text click operation, the clicked text information corresponding to the trigger operation in the operation image is identified; When the operation type is an icon click operation, extract the click icon image corresponding to the trigger operation from the operation image.

5. The method according to claim 1, characterized in that, The method further includes: Displays the chapter information entry interface; In response to the user's completion of input on the chapter information input interface, a chapter image containing chapter information is obtained; Identify the chapter information contained in the chapter image to obtain the target translated content corresponding to the chapter image; The step of generating the operation manual for the target application based on the target translated content includes: Based on the target translation content corresponding to the operation image and the target translation content corresponding to the chapter image, an operation manual for the target application is generated.

6. An apparatus for generating an operation manual, characterized in that, The device includes: The monitoring module is used to monitor user trigger operations in the display interface of the target application, determine the trigger information corresponding to the trigger operation from the current interface information of the display interface, the trigger information includes at least one of trigger object location information and trigger operation location information, and add image markers to the current interface information of the display interface based on the trigger information to obtain the operation image corresponding to the trigger operation; wherein, the target application refers to the application for which the operation manual is to be generated; The second generation module is used to generate image index information of the operation image corresponding to the triggering operation based on the triggering order of the triggering operation; The classification module is used to classify the operation image using an image classification algorithm to obtain the operation type corresponding to the operation image; The translation module is used to translate the operation image using a translation strategy that matches the operation type, so as to obtain the target translated content corresponding to the operation image; The first generation module is used to generate an operation manual for the target application based on the target translated content; wherein, when there are multiple operation images, the target translated content corresponding to each operation image is sorted according to the image index information of each operation image, and the operation manual for the target application is generated based on the sorted target translated content.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Claim settlement operation guiding method and device, computer equipment and storage medium

    CN114549220A

  • Equipment control method and device based on electronic specification, equipment and storage medium

    CN115562086A