Image processing method and device applied to intelligent agent
By combining the server and interactive end of the intelligent agent, the editing time sequence data of the image editing model is read, the editing status nodes are identified, and status prompts and thumbnails are generated. This solves the problem of low efficiency in image version management after multiple edits and improves the intuitiveness and convenience of image processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-03-27
AI Technical Summary
In the process of image processing, users have increasingly higher requirements for the intuitiveness and convenience of image processing, but existing technologies are unable to effectively manage and display multiple edited versions of images, resulting in low screening efficiency.
By coordinating the server and the interactive end of the intelligent agent, the editing time sequence data of the image editing model is read, the editing status nodes are identified and status prompts are generated, and the thumbnail of the snapshot image and status prompts are generated based on the editing detection results, so as to realize the synchronous display of the snapshot.
It improves the intuitiveness and convenience of image processing, reduces the time spent filtering invalid edited images, and enhances the user experience.
Smart Images

Figure CN121746534A_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of image processing technology, and in particular to an image processing method and apparatus for use with intelligent agents. Background Technology
[0002] With the continuous development of digital image processing and artificial intelligence technologies, AI-generated image applications and image processing tools have been widely deployed and used on various terminal devices and application platforms. Users can generate images and perform diverse processing on image content through these applications. However, in practical applications, users often modify or regenerate images multiple times to obtain images that meet their needs. In this case, as the number of image edits increases, the number of image versions gradually increases, and users' requirements for the intuitiveness and convenience of image processing continue to rise, how to better achieve image processing has become a key focus for all parties. Summary of the Invention
[0003] This specification provides one or more embodiments of an image processing method applied to an intelligent agent, comprising: reading editing time-series data of an image editing model called by the intelligent agent's server to edit a generated image; performing editing state recognition on the editing time-series data to obtain editing state nodes, and generating status prompts for the editing state nodes; performing image editing detection on a snapshot image based on the editing time-series data, and generating a thumbnail of the snapshot image according to the editing detection results; and synchronizing the thumbnail and the status prompts to the intelligent agent's interactive terminal for snapshot synchronization display.
[0004] This specification provides one or more embodiments of another image processing method applied to an intelligent agent, comprising: editing a generated image through cooperation between the intelligent agent's interactive terminal and a server. The method includes receiving a thumbnail and status prompt text of a snapshot image synchronized by the server from an edit status node; the edit status node is obtained by performing edit status recognition on the edit time sequence data of the generated image, and the thumbnail is generated based on the edit detection result obtained by performing image edit detection on the snapshot image. The snapshot is then displayed synchronously based on the thumbnail and the status prompt text.
[0005] This specification provides one or more embodiments of an image processing apparatus for an intelligent agent, comprising: a data reading module configured to read editing time-series data of an image editing model invoked by the intelligent agent's server to perform image editing on a generated image; a node recognition module configured to perform editing state recognition on the editing time-series data to obtain editing state nodes and generate status prompts for the editing state nodes; a thumbnail generation module configured to perform image editing detection on a snapshot image based on the editing time-series data and generate a thumbnail of the snapshot image according to the editing detection results; and a data synchronization module configured to synchronize the thumbnail and the status prompts to the interactive terminal of the intelligent agent for snapshot synchronization display.
[0006] This specification provides one or more embodiments of another image processing apparatus applied to an intelligent agent, comprising: an image editing module configured to perform image editing on a generated image in cooperation with a server via an interactive terminal of the intelligent agent; a data receiving module configured to receive a thumbnail and status prompt text of a snapshot image of an edit status node synchronized by the server; the edit status node is obtained by performing edit status recognition on the edit time sequence data of the generated image, and the thumbnail is generated based on the edit detection result obtained by performing image edit detection on the snapshot image; and a snapshot display module configured to synchronously display a snapshot based on the thumbnail and the status prompt text.
[0007] This specification provides one or more embodiments of an image processing device for an intelligent agent, comprising: a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to: read editing time-series data of an image editing model invoked by the intelligent agent's server to edit a generated image; perform editing state recognition on the editing time-series data to obtain editing state nodes, and generate status prompts for the editing state nodes; perform image editing detection on a snapshot image based on the editing time-series data, and generate a thumbnail of the snapshot image based on the editing detection results; and synchronize the thumbnail and the status prompts to the intelligent agent's interactive terminal for snapshot synchronization display.
[0008] This specification provides one or more embodiments of another image processing device for use with an intelligent agent, comprising: a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to: perform image editing on a generated image in cooperation with a server through an interactive terminal of the intelligent agent; receive a thumbnail and status prompt of a snapshot image synchronized by the server; the edit status node is obtained by performing edit status recognition on edit timing data of the generated image, and the thumbnail is generated based on the edit detection result obtained by performing image edit detection on the snapshot image; and perform snapshot synchronization display based on the thumbnail and the status prompt.
[0009] This specification provides one or more embodiments of a computer-readable storage medium for storing computer-executable instructions. When executed, these instructions implement the following process: reading editing timeline data of an image editing model invoked by a server of an intelligent agent to edit a generated image; performing editing state recognition on the editing timeline data to obtain editing state nodes and generating status prompts for the editing state nodes; performing image editing detection on a snapshot image based on the editing timeline data and generating a thumbnail of the snapshot image according to the editing detection results; and synchronizing the thumbnail and the status prompts to the interactive terminal of the intelligent agent for snapshot synchronization display.
[0010] This specification provides one or more embodiments of another computer-readable storage medium for storing computer-executable instructions, which, when executed, implement the following process: Image editing of a generated image is performed through cooperation between an intelligent agent's interactive terminal and a server. The process includes receiving a thumbnail and status prompt of a snapshot image synchronized by the server; the edit status node is obtained by performing edit status recognition on the edit timing data of the generated image, and the thumbnail is generated based on the edit detection result obtained by performing image edit detection on the snapshot image. A snapshot is then synchronously displayed based on the thumbnail and the status prompt. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Figure 1 A schematic diagram illustrating an implementation environment for an image processing method applied to an intelligent agent, provided in one or more embodiments of this specification; Figure 2A flowchart illustrating an image processing method applied to an intelligent agent, provided for one or more embodiments of this specification; Figure 3 A schematic diagram of a first image editing interface provided for one or more embodiments of this specification; Figure 4 A schematic diagram of a second image editing interface provided for one or more embodiments of this specification; Figure 5 A timing diagram of an image processing method applied to an intelligent agent in an image processing scenario, provided by one or more embodiments of this specification; Figure 6 A flowchart illustrating another image processing method applied to an intelligent agent, provided in one or more embodiments of this specification; Figure 7 A schematic diagram illustrating an embodiment of an image processing apparatus applied to an intelligent agent, provided by one or more embodiments of this specification. Figure 8 A schematic diagram of another embodiment of an image processing apparatus applied to an intelligent agent, provided by one or more embodiments of this specification; Figure 9 A schematic diagram of the structure of an image processing device applied to an intelligent agent, provided for one or more embodiments of this specification; Figure 10 This is a schematic diagram of another image processing device applied to an intelligent agent, provided for one or more embodiments of this specification. Detailed Implementation
[0012] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this document.
[0013] The image processing method for intelligent agents provided in one or more embodiments of this specification is applicable to the implementation environment of intelligent agents. (Refer to...) Figure 1 The implementation environment includes at least: The server of the intelligent agent is 101, the interaction terminal of the intelligent agent is 102, and the image editing model is 103; The server 101 of the intelligent agent is used to cooperate with the interaction terminal 102 of the intelligent agent to perform image editing on the generated image, and to read the editing time sequence data of the image editing model 103 in performing image editing on the generated image, and to perform image processing based on the editing time sequence data to obtain the thumbnail and status prompt words of the snapshot image corresponding to the editing status node, and to synchronize the thumbnail and status prompt words to the interaction terminal 102 of the intelligent agent; the server 101 of the intelligent agent can be deployed on a server, which can be one or more servers, a server cluster composed of several servers, or a cloud server of a cloud computing platform. The intelligent agent's interaction terminal 102 is used to cooperate with the server 101 to edit the generated image, and to receive thumbnails and status prompts of the snapshot images of the edit status nodes synchronized by the server 101, and to display the snapshots synchronously based on the thumbnails and status prompts; the intelligent agent's interaction terminal 102 can be deployed on terminal devices, which can be mobile phones, personal computers, tablets, e-book readers, wearable devices, devices that interact with information based on AR (Augmented Reality) / VR (Virtual Reality), and laptop computers, etc. The image editing model 103 is used to perform image editing on the generated image based on the editing instructions for the generated image in response to the call of the server 101 of the agent. It can also be used to store the editing timing data of the image editing, and to perform image processing on the generated image in response to the call of the server 101, and output status prompts and thumbnails.
[0014] In addition, the implementation environment may also include a snapshot generation model 104, which can perform snapshot generation processing in response to a call from the image editing model 103.
[0015] In this implementation environment, during image processing, the server 101 of the intelligent agent can cooperate with the interactive terminal 102 of the intelligent agent to call the image editing model 103 to edit the generated image and obtain editing time sequence data. Based on this, the server 101 first reads the editing time sequence data of the image editing model 103 that edits the generated image, performs editing state recognition on the editing time sequence data to obtain editing state nodes, and generates status prompt words for the editing state nodes. Further, based on the editing time sequence data, it performs image editing detection on the snapshot image corresponding to the editing state node, generates a thumbnail of the snapshot image according to the editing result, and then synchronizes the thumbnail and status prompt words to the interactive terminal 102 of the intelligent agent. Correspondingly, the interactive terminal 102 receives the thumbnail and status prompt words of the snapshot image of the editing state node synchronized by the server, and performs snapshot synchronization display based on the thumbnail and status prompt words. In this way, image processing of the generated image is achieved by calling the image editing model.
[0016] This specification provides one or more embodiments of an image processing method applied to intelligent agents, as follows: Reference Figure 2 The image processing method for intelligent agents provided in this embodiment can be applied to the server side of intelligent agents. The method specifically includes steps S202 to S208.
[0017] Step S202: Read the editing timing data of the image editing model called by the agent's server to perform image editing on the generated image.
[0018] In this embodiment, an intelligent agent refers to an executor capable of autonomously performing tasks, making decisions, and learning and adjusting according to environmental changes. Specifically, an intelligent agent can be an executor integrating image generation and image editing capabilities. The intelligent agent can receive natural language commands or interactive operations input by the user, invoke one or more models, drive the models to perform corresponding image processing, and manage state data during the image editing process. Here, an intelligent agent can specifically be an intelligent agent, an intelligent agent application, an intelligent agent system, or a large language model (LLM), such as a large language model using the Transformer architecture. Alternatively, an intelligent agent can also be an image generation model, such as an image generation model using a neural network architecture.
[0019] The image editing model refers to a model used to edit the generated image, specifically a model for image editing of the generated image. The input to the image editing model may include the generated image and / or image editing instructions input by the user. Here, when the aforementioned agent is an image generation model, the image editing model and the image generation model can be one model or two models; this embodiment does not limit this. Optionally, the generated image is obtained by the agent performing image generation.
[0020] The editing time sequence data refers to the set of data related to image editing recorded in chronological order or in the order of operation during the editing operation of the generated image; the editing time sequence data may include the edited image, editing archive instruction, editing time and / or editing parameters corresponding to each editing step.
[0021] In practical implementation, the server-side of the intelligent agent can cooperate with the client-side to generate and obtain the generated image. For example, after receiving text prompts submitted by the client-side, the server-side can generate an image based on the text prompts using the intelligent agent. Based on this, once the generated image is obtained, the server-side can cooperate with the client-side to edit it. Specifically, when the server receives an editing command for the generated image initiated by the user through the client-side, it can call an image editing model to perform corresponding image editing processing. During this process, editing timeline data can be generated based on the editing data of each image editing process performed on the generated image. Subsequently, during image processing, to ensure data integrity and comprehensiveness, the editing timeline data is read, that is, the editing timeline data of the image editing model used to edit the generated image is read. Optionally, the image editing model can be called after detecting an editing command for the generated image.
[0022] It should be noted that during the process of image editing between the server and the client, the server can synchronize the edited image obtained from the image editing process to the client in real time.
[0023] Step S204: Perform editing status recognition on the editing time sequence data to obtain editing status nodes, and generate status prompt words for the editing status nodes.
[0024] In practice, after image editing is performed on the generated image based on editing instructions, considering that image editing may produce a large number of edited images, in order to avoid low screening efficiency due to an excessive number of invalid edited images, the editing time series data is used to identify the editing state nodes. On this basis, in order to convert the relevant editing parameters and / or image feature modifications in the editing time series data into semantic information, status prompt words for the editing state nodes are further generated.
[0025] The edit state node refers to a node that can represent a specific editing stage or editing state; specifically, it refers to a snapshot node in the editing time series data that can be used to represent a specific state during the image editing process; for example, the edit state node can be a core object change node in the generated image, or an image style change node in the generated image; here, the edit state node can also be replaced with a snapshot node, and correspondingly, the following description of the edit state node can also be replaced with a snapshot node.
[0026] The status prompt words refer to text used to describe a specific editing state, operation content, and / or core information; specifically, they refer to editing prompt words or editing keywords generated based on the editing operation, image changes, and / or editing parameters corresponding to the editing state node. Status prompt words can be used to describe the key or major editing changes of the image in the current editing state node compared to the previous editing state node, so as to inform the user of the core editing content of the current editing state node; here, status prompt words can also be replaced with editing prompt words, editing keywords, or status keywords, and correspondingly, the following descriptions of status prompt words can also be replaced with editing prompt words, editing keywords, or status keywords.
[0027] In the specific execution process, during the editing state identification of editing time-series data, in order to improve the comprehensiveness of editing state identification, on the one hand, editing difference nodes can be obtained through editing difference detection, and on the other hand, editing key nodes can be obtained through similarity sequence analysis. In one optional implementation method provided in this embodiment, the editing state identification of editing time-series data to obtain editing state nodes includes: Edit difference detection is performed on the edited images contained in the editing time series data to obtain edit difference nodes; The similarity sequence is obtained by performing similarity calculation on the edited images contained in the editing time series data, and the editing key nodes are obtained by performing editing key identification based on the similarity sequence.
[0028] Specifically, on the one hand, in the process of determining editing difference nodes, editing difference detection can be performed on the edited image to obtain editing difference nodes. For example, visual algorithms and image feature analysis can be used to compare edited images at different editing stages in the editing time series data to identify changes in image content, image features, and / or image parameters, so as to determine the image nodes that produce significant editing effects. Here, editing difference detection can also be replaced by editing saliency detection, or it can also be replaced by editing saliency detection. On the other hand, in the process of determining the key editing nodes, similarity calculations can be performed on the features of adjacent or related edited images in the editing time series data, and a quantitative similarity index can be output to reflect the degree of similarity between different edited images. A similarity sequence is obtained based on each similarity index, and the key editing nodes are obtained by identifying the key editing nodes through the similarity sequence. Here, the key editing identification can also be replaced by the focus editing identification, and the key editing nodes can also be replaced by the focus editing nodes.
[0029] In the specific execution process, during the editing state identification of the editing time-series data, in order to improve the accuracy and objectivity of the editing state identification, feature quantification can be performed based on the visual and / or semantic features of each edited image. Furthermore, by calculating the feature difference value between the index image and the index subject image, the editing changes between the edited images are transformed into quantifiable numerical indicators, thereby determining the editing state nodes based on these numerical indicators. In another optional implementation of this embodiment, obtaining editing state nodes by identifying the editing state of the editing time-series data includes: Visual and semantic features are extracted and merged from each edited image in the editing time series data to obtain visual and semantic features; Calculate the feature difference value between the visual semantic features of the index image and the visual semantic features of the main index image in each edited image, and determine the edit state node based on the obtained feature difference value.
[0030] Here, the index image refers to the edited image that is currently being evaluated or is currently to be analyzed; in the process of identifying the edit status of the edit time series data, each frame of the edit time series data can be used as an index image to participate in the feature difference calculation during the traversal of the edit time series data. The index subject image refers to the reference image used as the benchmark for feature difference calculation, specifically the reference image used for comparison with the index image. The index subject image can be the initially generated image, or it can be the edit image corresponding to the previously determined edit state node in the edit time series data. The index subject image can also be replaced with the index target image, or it can also be replaced with the index object image. Correspondingly, the following description of the index subject image can also be replaced with the index target image or the index object image.
[0031] Specifically, in the process of editing state recognition of editing time-series data, visual and semantic features can be extracted from each editing image in the editing time-series data. Then, feature fusion is performed on the obtained visual and semantic features of each editing image to obtain the visual and semantic features of each editing image. Based on this, in the process of calculating the feature difference between the index image and the index subject image, the initial generated image in the editing image sequence can be used as the index subject image. The visual and semantic feature vector of this index subject image is extracted as the baseline feature vector. Then, the next editing image is selected sequentially as the index image, and its visual and semantic feature vector is extracted. The differences between the index image and the index subject image are then calculated. If the feature difference value between the visual semantic feature vectors is greater than the difference determination threshold, the node corresponding to the index image is determined to be an edit state node, and the index image is updated to a new index subject image. Conversely, if the feature difference value of the current index image is less than or equal to the difference determination threshold, the node corresponding to the index image is determined to be a non-critical node and is not included in the edit state node set, and the index subject image remains unchanged. This process of feature difference calculation is repeated for all edit images in the edit image sequence in chronological order until the difference calculation and edit state node determination of all edit images are completed, and finally the edit state node set is obtained.
[0032] It should be added that, in the process of obtaining edit state nodes by identifying the edit state of the edit time series data, in addition to the visual and semantic features obtained by extracting and merging visual and semantic features from the edited image as described above, multi-dimensional feature extraction can also be performed on the edited image to obtain multi-dimensional image features, and feature difference calculation can be performed based on the multi-dimensional image features. Based on this, the implementation method provided above can also be replaced as follows: multi-dimensional feature extraction is performed on the edited images contained in the edit time series data to obtain multi-dimensional image features; the feature difference score between the multi-dimensional image features of the index image and the multi-dimensional image features of the index object image in each edited image is calculated; the edit state image is determined based on the obtained feature difference score, and the node corresponding to the edit state image is taken as the edit state node. Here, the process of calculating the feature difference based on the multi-dimensional image features is similar to the process of calculating the feature difference based on the visual and semantic features provided above. You can refer to the process of obtaining edit state nodes by identifying the edit state of the edit time series data as described above and make adaptive modifications, changes or deletions according to the needs of the actual execution process.
[0033] In this process of determining the edit state node based on feature difference values, in order to adapt to the difference distribution patterns of different image editing processes and / or different types of edited images, an adaptive threshold can be introduced, and the edit state node can be determined based on the feature difference values and the adaptive threshold. In one optional implementation of this embodiment, determining the edit state node based on the obtained feature difference values includes: The significance of differences between each edited image is calculated based on the feature difference values. Edit nodes corresponding to edited images with significantly different values greater than the adaptive threshold are identified as edit state nodes.
[0034] The adaptive threshold refers to a threshold that is dynamically adjusted based on the input data. Specifically, it refers to a judgment threshold that is dynamically generated based on the numerical distribution characteristics of the significant differences in the values of each edited image. The adaptive threshold can be automatically adjusted according to the difference distribution patterns of different image editing processes and / or different types of edited images, rather than a fixed preset value. Optionally, the adaptive threshold is determined based on the numerical distribution index of the significant differences in the values of each edited image.
[0035] Specifically, once feature difference values are obtained, the significance of the difference can be calculated based on these values. For example, the feature difference values can be standardized, normalized, and / or weighted to convert them into significant difference values. Then, an adaptive threshold is determined based on each significant difference value. Once the adaptive threshold is determined, the significant difference values are compared with the adaptive threshold, and the edit nodes corresponding to the edited images with significant difference values greater than the adaptive threshold are identified as edit status nodes.
[0036] In addition, during the specific execution process, in the process of identifying the editing status of the editing time sequence data, it is also possible to identify the user's specific editing mode in the image editing process, thereby locating the user's core decision nodes in the image editing process, so that the editing status nodes obtained by the editing status identification are more in line with the user's actual image editing intentions; In one optional implementation of this embodiment, editing status nodes are obtained by identifying the editing status of editing time-series data, including: Feature similarity sequences are obtained by performing feature similarity calculations on adjacent edited images contained in the editing time series data; The system detects similarity subsequences with specific editing patterns in the feature similarity sequence, verifies the editing patterns of these subsequences, and determines the editing state nodes in the verified similarity subsequences.
[0037] In this context, a specific editing mode refers to a recurring editing behavior and / or editing type that reflects user intent during image editing. For example, it could be a hesitation pattern exhibited by the user during multiple rounds of image editing, such as "repeated modification behavior" or "repeated redo behavior." Here, a specific editing mode can also be replaced with a specific editing behavior or a specific editing type. Correspondingly, the following descriptions of specific editing modes can also be replaced with specific editing behaviors or specific editing types.
[0038] It should be noted that the above-mentioned implementation methods for obtaining edit status nodes by identifying the edit status of edit time-series data can be implemented in any way according to actual needs. Alternatively, any two or all three implementation methods can be combined according to the actual execution process. Or, any two or all three implementation methods can be combined after adaptive modification, alteration, or deletion according to the actual execution process. For example, obtaining edit status nodes by identifying the edit status of edit time-series data includes: performing similarity calculation on the edited images contained in the edit time-series data to obtain a similarity sequence, and identifying edit key nodes based on the similarity sequence; extracting and merging visual and semantic features of each edited image contained in the edit time-series data to obtain visual semantic features; calculating the feature difference value between the visual semantic features of the index image and the visual semantic features of each index subject image in each edited image, and determining the edit status node based on the obtained feature difference value. For example, obtaining edit state nodes by identifying edit state in editing time-series data includes: extracting multidimensional features from the edited images contained in the editing time-series data to obtain multidimensional image features; calculating the feature difference score between the multidimensional image features of the index image and the multidimensional image features of the index target image in each edited image; determining the edit state image based on the obtained feature difference score, and taking the node corresponding to the edit state image as the edit state node; calculating the feature similarity of adjacent edited images contained in the editing time-series data to obtain a feature similarity sequence; detecting the existence of similarity subsequences with specific editing patterns in the feature similarity sequence, and verifying the editing pattern of the specific editing pattern in the similarity subsequences, and determining the edit state node in the similarity subsequences that pass the verification.
[0039] In practical applications, when determining edit status nodes, considering the diversity of users' image editing behaviors, the various edit status nodes determined by the above-mentioned edit status recognition algorithms may deviate from the user's operation. For example, the user may not recognize the important fine-tuning version as an edit status node by the algorithm. Therefore, edit status nodes can also be determined based on the user's editing and archiving instructions, and / or based on the configured editing time and / or number of edits. In one optional implementation of this embodiment, after reading the editing timing data of the image editing model called by the agent's server to edit the generated image, the following operations can also be performed: The editing status node is determined based on the editing archive instructions contained in the editing timing data, and / or, the editing status node is determined based on the editing time and editing status cycle contained in the editing timing data.
[0040] Among them, the edit archiving command refers to the command actively triggered by the user through the intelligent agent interaction terminal during the image editing process to save the current image editing version and / or image editing state; for example, the user can submit the edit archiving command by clicking the "Save Version" button, or through a preset shortcut key, or through voice command, etc. Figure 3 As shown, users can submit editing and archiving instructions by clicking the "Save Project" button 301.
[0041] The editing state cycle refers to a pre-configured, regular interval, specifically a pre-configured cycle for automatically triggering the generation of editing state nodes. The editing state cycle can include an editing time cycle and / or an editing operation cycle. The editing time cycle, for example, is a fixed time interval of 'a' minutes, and the editing operation cycle, for example, is every 'b' image editing operations. For instance, a user can click... Figure 3 The "Auto-save Settings" control 302 shown is used to configure the editing state cycle.
[0042] Specifically, in the process of determining the edit status node based on the edit archiving instruction, it can be detected whether the edit timing data contains the edit archiving instruction issued by the user. If it exists, the edit node corresponding to the specific image editing operation in the edit timing data corresponding to the edit archiving instruction is determined as the edit status node. In the process of determining the edit status node according to the edit time and / or the edit status cycle, when the accumulated time since the previous edit status node reaches the set edit time cycle, or the accumulated number of edit operations since the previous edit status node reaches the set edit operation cycle, the position of the edited image corresponding to the current edit operation in the edit timing data is determined as the edit status node.
[0043] In specific implementation, after determining the edit status node as described above, status prompt words for the edit status node can be further generated. Specifically, in the process of generating status prompt words, in order to ensure that the status prompt words can accurately reflect the core content of the edit, identify the content being edited, and generate corresponding status prompt words, in an optional implementation provided in this embodiment, generating status prompt words for the edit status node includes: Based on the image features of the snapshot image and the image features of the preceding adjacent snapshot image, feature difference calculation and difference localization are performed to obtain the editing difference region; Image semantic recognition is performed based on the edit difference region to obtain the edit difference object and / or edit difference type, and edit prompt words are generated based on the edit difference object and / or edit difference type. Specifically, the snapshot image refers to the snapshot image corresponding to the edit state node.
[0044] The edit difference region refers to the sub-region of the edited image obtained after difference localization. Specifically, it refers to the sub-region where the image content of the current snapshot image has changed compared with the previous adjacent snapshot image, such as the sub-region where the image content has been modified, added, and / or deleted.
[0045] The editable difference object refers to an object whose content changes between two or more images. Specifically, it refers to an object that is added, deleted, modified, and / or replaced during the image editing process compared to the previous adjacent snapshot image. For example, the editable difference image can be a specific physical entity, such as a person or building, or it can be a visual component, such as a background or texture area.
[0046] The edit difference type refers to the category of editing behavior determined based on the nature or change pattern of image editing operations, such as add type, delete type, replace type, and modify type. In addition, it can also be other editing behavior categories, such as style transfer type and attribute adjustment type.
[0047] Specifically, given the aforementioned determination of edit state nodes, the snapshot image corresponding to each edit state node is obtained. Based on this, the current snapshot image is compared with its preceding adjacent snapshot image. Specifically, image features are extracted from the current snapshot image and the preceding adjacent snapshot image respectively, and feature differences are calculated for the image features of the two snapshot images. The two snapshot images are aligned, and a region-by-region comparison is performed based on the image features of the two snapshot images. The degree of feature difference between the two snapshot images in the corresponding regions is calculated, and regions with a feature difference degree exceeding a preset threshold are marked as editing difference regions. Based on the identified editing difference region, the editing difference region can be cropped and extracted from the original snapshot image, and image semantic recognition can be performed to obtain the editing difference object and / or editing difference type. Core elements are extracted for the editing difference object and / or editing difference type, and prompt words are generated to obtain editing prompt words corresponding to the editing difference object and / or editing difference type.
[0048] For example, feature difference calculation and difference localization are performed based on the image features of a snapshot image and the image features of the preceding adjacent snapshot image. The obtained editing difference region corresponds to the region where the car is deleted. Image semantic recognition is performed based on the editing difference region to obtain the editing difference object as "car" and the editing difference type as "delete". Then, the editing prompt word generated based on the editing difference object and the editing difference type is "delete car".
[0049] Furthermore, considering that during image editing, users may not perform substantial additions, deletions, or replacements of core editing objects and / or semantic content types, but rather fine-tune editing parameters for the generated image, in order to comprehensively cover various editing operations and ensure that each editing state node generates accurate status prompts, status prompts for the editing state nodes can also be generated based on the editing parameters to provide editing parameter prompts. In another optional implementation provided in this embodiment, generating status prompts for editing state nodes includes: Read the editing parameters of each edited image in the editing time segment corresponding to the editing status node, and perform parameter fusion on the read editing parameters to obtain editing parameter prompts.
[0050] Among them, editing parameters refer to the parameters used to adjust the visual effects of the generated image. Specifically, they can be visual adjustment parameters or visual editing parameters. For example, editing parameters can be brightness parameters, contrast parameters, saturation parameters and / or sharpness parameters. In addition, they can also be other parameters used to adjust the visual effects of the image, such as hue parameters.
[0051] An editing time sequence refers to a continuous editing operation time interval corresponding to a single editing state node. Specifically, it refers to the editing subsequence consisting of all edited images and corresponding editing time sequence data from the previous adjacent editing state node to the current editing state node.
[0052] Specifically, given the aforementioned determination of the edit status node, the editing parameters of each edited image in the edit time segment corresponding to the edit status node are read from the edit time sequence data. The read editing parameters are then merged, deduplicated, and / or weighted and summarized to obtain the editing parameter prompt words.
[0053] For example, if a user adjusts the editing parameters of a generated image multiple times, specifically adjusting the contrast, the editing parameters of each image in the editing time segment corresponding to the editing state node can be read, and the editing parameter prompt obtained by fusing the read editing parameters can be "adjust contrast".
[0054] It should be noted that the above-mentioned implementation methods for generating status prompts for edit status nodes can be implemented in any way according to actual needs. Alternatively, all two implementation methods can be combined according to the needs of the actual execution process. Or, all two implementation methods can be combined after adaptive modification, alteration, or deletion according to the needs of the actual execution process. For example, generating status prompts for edit status nodes includes: calculating and locating the feature differences between the image features of the snapshot image and the image features of the preceding adjacent snapshot images to obtain the edit difference region; performing image semantic recognition based on the edit difference region to obtain the edit difference object, and generating an edit prompt based on the edit difference object; reading the editing parameters of each edited image in the editing time segment corresponding to the edit status node, and performing parameter fusion on the read editing parameters to obtain the edit parameter prompt; here, the obtained edit prompt and edit parameter prompt are the status prompts for the edit status node.
[0055] Step S206: Perform image editing detection on the snapshot image based on the editing time sequence data, and generate a thumbnail of the snapshot image based on the editing detection results.
[0056] In specific implementation, given the aforementioned determination of edit state nodes, snapshot images corresponding to each edit state node can be further generated. Based on this, in order to accurately characterize the core editing features of the snapshot image compared to the previous snapshot image, and to transform the edit state nodes into visual elements to support the visualization of the edit state nodes, image editing detection is performed on the snapshot images based on the editing time-series data. Thumbnails of the snapshot images are generated based on the editing detection results. Here, the generated thumbnails of the snapshot images are also known as snapshot thumbnails. Based on this, thumbnails can also be replaced with snapshot thumbnails. Correspondingly, the following descriptions of thumbnails can also be replaced with snapshot thumbnails.
[0057] In this context, a thumbnail refers to a preview image generated through image processing based on the core editing features of a snapshot image. It is a low-resolution and / or small-size representation of the snapshot image. It's important to note that a thumbnail is not a uniform compression of the entire snapshot image; rather, it is generated through image processing based on a cropping and focusing strategy determined by the editing detection results. Figure 1 On the one hand, it can preserve the main visual content in the image, such as the object being edited; on the other hand, it can prioritize ensuring that the areas of difference in editing are visible and prominent in the thumbnail. For example, if a user only adds a car in the lower right corner of the image, even if the area does not occupy the center position in the original generated image, the thumbnail will retain the added car in the thumbnail through cropping and / or layout adjustments, avoiding the loss of key editing information due to regular scaling.
[0058] In the specific execution process, during the generation of thumbnails for snapshot images, in order to improve the core feature recognition of the generated thumbnails, the core content of the snapshot image can be preserved while highlighting the core changes in the editing process. Based on this, thumbnails for snapshot images can be generated based on the regions of interest and / or regions of editing difference in the snapshot image. In one optional implementation of this embodiment, image editing detection is performed on the snapshot image based on editing time-series data, and thumbnails for the snapshot image are generated based on the editing detection results, including: Regions of interest are obtained by performing region of interest detection on the snapshot image, and edit difference recognition is performed on the snapshot image to obtain the edit difference region; The thumbnail generation parameters are determined based on the region of interest and the region of difference in editing, and the snapshot image is generated into a thumbnail according to the thumbnail generation parameters.
[0059] The region of interest refers to the core content area in the snapshot image, specifically the core area with visual salience in the snapshot image. For example, the region of interest can be the area containing the key visual elements of the snapshot image.
[0060] In addition to generating thumbnails based on the thumbnail generation parameters determined according to the region of interest and the region of editing difference, as mentioned above, considering that in some editing scenarios, it is difficult to intuitively distinguish the type of editing operation simply by cropping and / or scaling, in order to further improve the thumbnail's ability to identify the edited content and to make different types of editing operations clearly visually distinguishable in the thumbnail, the editing object and / or editing area of the snapshot image can also be explicitly marked to obtain the thumbnail; another optional implementation provided in this embodiment is to perform image editing detection on the snapshot image based on editing time-series data, and generate a thumbnail of the snapshot image based on the editing detection results, including: The system performs object recognition and / or region detection on the snapshot image, and performs image segmentation on the obtained object and / or region to obtain the target image block. Thumbnails are obtained by explicitly marking the target image blocks according to the explicit marking method corresponding to the editing type of the editing object and / or editing area.
[0061] The explicit marking method refers to the pre-defined and distinctive marking methods for different editing types, such as color highlighting, outline indication, texture filling and / or transparency adjustment, to highlight the areas of editing difference, so that the editing changes are intuitively distinguishable in the thumbnail.
[0062] Specifically, in the process of generating thumbnails for snapshot images, for snapshot images with content editing, the editable objects can be identified to obtain the editable objects; for snapshot images with parameter adjustments, the editable regions can be detected to obtain the editable regions. After identifying the editable objects and / or detecting the editable regions, image segmentation is further performed on the identified editable objects and / or detected editable regions to obtain target image blocks. Subsequently, the corresponding explicit marking method is determined according to the editing type of the editable objects and / or editable regions, and the target image blocks are explicitly marked to obtain thumbnails.
[0063] Here, the explicit marking method can be determined according to the editing type of the editing object and / or editing area, or it can be determined according to the target image block, or it can be determined according to the editing parameters corresponding to the target image block. Based on this, the explicit marking method corresponding to the editing type of the editing object and / or editing area can be replaced with: the explicit marking method corresponding to the target image block, or it can be replaced with: the explicit marking method corresponding to the editing parameters corresponding to the target image block, or it can be replaced with: determining the explicit marking method corresponding to the target image block and following the explicit marking method corresponding to the target image block.
[0064] Step S208: Synchronize the thumbnail and the status prompt to the interactive terminal of the intelligent agent to perform snapshot synchronization display.
[0065] In practice, once the thumbnail and status prompt are obtained, they are synchronized to the agent's interactive terminal. Subsequently, the agent's interactive terminal receives the thumbnail and status prompt of the snapshot image of the edited status node and displays the snapshot synchronously based on the thumbnail and status prompt.
[0066] For example, after the agent's server synchronizes thumbnails and status prompts to the agent's interface, the interface can perform a snapshot synchronization display; specifically, during the snapshot synchronization display process on the interface, the interface can... Figure 3 The image editing interface shown displays thumbnails, such as thumbnails 303-1, 303-2 and 303-3, as well as status prompts, such as status prompt 304-1 corresponding to thumbnail 303-1; In addition, the image editing interface displayed on the intelligent agent's interactive terminal can also display image editing and processing controls and an "image export" control. Image editing and processing controls include controls such as "stylize," "background blur," "intelligent crop," and "remove object." Furthermore, if the user triggers any thumbnail in thumbnail 303, such as clicking thumbnail 303-3, the snapshot details of the snapshot image corresponding to thumbnail 303-3 can be further displayed in the image editing interface. Specifically, the snapshot details can display basic information of the snapshot image (such as operation type, creation time), editing parameters, and / or change records. If the user triggers the "delete control" 305 in the snapshot details, the current snapshot image can be deleted.
[0067] In the specific execution process, during the snapshot synchronization display on the interactive end, in order to eliminate the comparison error caused by image offset, improve the visualization of differences in edited images, and improve the efficiency of users' snapshot image backtracking and comparison operations, the snapshot image corresponding to any thumbnail can be aligned with the preceding adjacent snapshot image, and the aligned snapshot image can be compared and displayed. In an optional implementation method provided in this embodiment, the snapshot synchronization display is implemented in the following way: The interactive image editing interface displays thumbnails and status prompts; If any thumbnail displayed in the image editing interface is triggered, the snapshot image corresponding to any thumbnail is aligned with the preceding adjacent snapshot image, and the aligned snapshot image is compared and displayed according to the layered display.
[0068] Specifically, after receiving the thumbnail and status prompt of the snapshot image of the edit status node synchronized by the server, the agent's interactive terminal displays the thumbnail and status prompt in the image editing interface of the interactive terminal. Subsequently, if any thumbnail displayed in the image editing interface is triggered, the snapshot image corresponding to the triggering thumbnail can be registered with the preceding adjacent snapshot image at the pixel level and / or feature level, so that the positions of the same background and main object in the two snapshot images are completely overlapped, that is, the target snapshot image corresponding to any triggered thumbnail is image aligned with the reference snapshot image. After that, based on the image alignment, two layers can be created in the editing canvas of the interactive terminal, and the image-aligned target snapshot image and the reference snapshot image can be overlaid and displayed as two independent layers.
[0069] For example, in Figure 3In the image editing interface shown, if thumbnail 303-3 is triggered, the snapshot image corresponding to thumbnail 303-3 is aligned with the preceding adjacent snapshot image. That is, the snapshot image corresponding to thumbnail 303-3 is aligned with the snapshot image corresponding to thumbnail 303-2. The image editing interface then compares and displays the aligned snapshot images according to layers, as shown below. Figure 4 As shown, Figure 4 The image editing interface shown displays snapshot image 401 corresponding to thumbnail 303-3 and snapshot image 402 corresponding to thumbnail 303-2. The background textures of snapshot image 401 and snapshot image 402 are completely overlapped. Users can adjust the screen ratio of snapshot image 401 and snapshot image 402 by dragging the "Drag to view differences" control 403. If the progress bar of the "Drag to view differences" control is slid to the left, the display ratio of snapshot image 401 corresponding to thumbnail 303-3 will gradually increase, while the display ratio of snapshot image 402 on the right will decrease accordingly. If the progress bar of the "Drag to view differences" control is slid to the right, the display ratio of snapshot image 402 will increase. Alternatively, users can also adjust the screen ratio of snapshot image 401 and snapshot image 402 by using control 404.
[0070] In this embodiment, the image processing implementation method for the intelligent agent provided above can also be performed by an image editing model. Specifically, when the image processing implementation method for the intelligent agent provided above is performed by an image editing model, the server of the intelligent agent can call the image editing model to enable the image editing model to perform editing state recognition on the editing time sequence data to obtain editing state nodes and generate status prompt words for the editing state nodes. Based on the editing time sequence data, image editing detection is performed on the snapshot image, and a thumbnail of the snapshot image is generated according to the editing detection result. Subsequently, the image editing model outputs status prompt words and thumbnails. Correspondingly, the thumbnail and status prompt words output by the image editing model are obtained, and the thumbnail and status prompt words are synchronized to the interactive end of the intelligent agent for snapshot synchronization display. In this process, specifically during image processing in the image editing model, the first step is to determine the editing state node. This involves first extracting visual features from each edited image in the editing time series data using a Convolutional Neural Network (CNN), and then extracting semantic features from each edited image using a Transformer with a self-attention mechanism. Visual semantic features are obtained by merging features through a feature concatenation and weighted fusion network. Then, the feature similarity between adjacent edited images is calculated using a cosine similarity calculator, and the feature difference value is calculated using an L2 distance calculator. Significant difference values are generated based on the ratio of the difference value to a preset baseline value. The similarity sequence is then time-series modeled using a Recurrent Neural Network (RNN) to detect similarity subsequences that conform to a specific editing mode. Validation is completed by comparing with a preset mode template, and the validated similarity subsequences are output. An adaptive threshold generator determines an adaptive threshold based on the mean and variance of the distribution of significant difference values, and the editing state node is output through a logic editor. Then, status prompt words are generated. First, a feature matching algorithm is used to perform point-by-point matching of the fused features of the two snapshot images, marking the coordinates of the difference regions and outputting the edit difference regions. Based on the edit difference regions, image semantic recognition is performed to obtain the edit difference objects and / or edit difference types. Then, the prompt word generation module maps the edit difference object labels and / or edit difference type labels into natural language expressions and outputs status prompt words. After that, thumbnail generation is performed. First, regions with high visual weight in the image are selected through multi-scale segmentation and similarity calculation to determine the regions of interest. Then, the edit difference regions are determined through feature matching algorithms and pixel-by-pixel comparison. The edit objects are detected by Mask R-CNN and the corresponding segmentation masks are generated. Then, the cropping range is adjusted and the appropriate scaling ratio is determined by the adaptive cropping strategy generator according to the regions of interest and edit difference regions. The labeling method is selected based on the edit type label. The target image patch is labeled and rendered by the image renderer. Finally, the image scaling and compression algorithm is used for cropping and scaling to output the thumbnail.
[0071] Furthermore, the image processing implementation method for intelligent agents provided above can also be performed by a snapshot generation model. Specifically, the snapshot generation model can perform snapshot generation processing in response to a call from the image editing model. Specifically, when the image processing implementation method for intelligent agents provided above is performed by a snapshot generation model, after the image editing model completes the image editing operation and generates editing time-series data, it can further call the snapshot generation model to perform snapshot generation processing. Specifically, during the snapshot generation process, the snapshot generation model performs editing state recognition on the editing time-series data to obtain editing state nodes and generates status prompts for the editing state nodes based on the editing time-series data. The data performs image editing detection on the snapshot image and generates a thumbnail of the snapshot image based on the editing detection results. Subsequently, after obtaining the thumbnail and status prompt, the snapshot generation model can push the result to the callback interface of the image editing model according to the callback configuration information in the call request. The image editing model, after obtaining the thumbnail and status prompt, further returns the thumbnail and status prompt to the agent's server. Correspondingly, the thumbnail and status prompt output by the image editing model are obtained and synchronized with the agent's interactive end for snapshot synchronization display. Optionally, the snapshot generation model responds to the call of the image editing model to perform snapshot generation processing.
[0072] It should be noted that the snapshot generation process performed by the snapshot generation model is similar to the image processing process performed by the image editing model provided above. Please refer to the specific process of image processing performed by the image editing model provided above. This embodiment will not repeat it here.
[0073] It should be added that each optional implementation method and each feasible execution method in steps S202 to S208 provided in this embodiment can be executed independently as needed, or they can be combined and referenced with each other. At the same time, each specific execution step in each optional implementation method or each feasible execution method can also be executed independently or combined as needed. Any feature in each execution step can also be deleted, or any feature in one execution step can be added to another execution step or replace any feature in another execution step. The execution conditions of "if" or "under what circumstances" involved in each step or operation can be directly deleted. This embodiment does not specifically limit the subsequent operations after the execution conditions.
[0074] It should also be added that, depending on the actual application scenario, step S202 and any of the subsequent steps S204 to S208 can be deleted, or any feature in any step can be deleted. For example, the reading and editing timing data in step S202 can be deleted, and the execution order of steps S202 to S208 can also be arbitrary.
[0075] In summary, the image processing method for intelligent agents provided in this embodiment, in order to ensure data integrity and comprehensiveness during image processing, firstly reads the editing time sequence data of the image editing model being called to edit the generated image. Based on this, considering the large number of edited versions generated by multiple rounds of image editing and the high data redundancy, to avoid low filtering efficiency due to an excessive number of invalid edited versions, the editing time sequence data is used to identify editing state nodes and generate status prompts for these nodes. Furthermore, snapshot images corresponding to each editing state node can be generated. Subsequently, to accurately characterize the snapshot image compared to the previous one... The core editing features of the preceding snapshot image are transformed into intuitive visual elements. Image editing detection is performed on the snapshot image based on editing time-series data. Thumbnails of the snapshot image are generated based on the editing detection results. Finally, the thumbnails and status prompts are synchronized to the interactive end of the agent so that the interactive end of the agent can receive the thumbnails and status prompts and display the snapshot synchronously based on the thumbnails and status prompts. In this way, by improving the accuracy of the representation and the intuitiveness of the editing features of each edited image, the simplified and structured management of each edited image is achieved, and the user's understanding of the version management of each edited image is reduced, thereby improving the efficiency of the user in reviewing the editing of each edited image.
[0076] Steps S202 to S208 provided in this embodiment can be executed by the server of the intelligent agent. It should be noted that the steps S202 to S208 executed by the server of the intelligent agent and steps S602 to S606 executed by the interactive terminal of the intelligent agent in the following embodiment can cooperate with each other during the execution process. Therefore, when reading this embodiment, please refer to the corresponding content of steps S602 to S606 provided in the following method embodiment, and when reading the following method embodiment, please refer to the corresponding content of steps S202 to S208 provided in this embodiment.
[0077] The following example uses the image processing method for intelligent agents provided in this embodiment as an example in an image processing scenario, combined with... Figure 5 The image processing method applied to intelligent agents provided in this embodiment will be further described below. Figure 5 An image processing method for intelligent agents, applicable to image processing scenarios, specifically includes the following steps.
[0078] Step S502: Read the editing timing data of the image editing model called by the agent's server to perform image editing on the generated image.
[0079] Step S504: Perform edit difference detection on the edited images contained in the edit time series data to obtain edit difference nodes.
[0080] Step S506: Perform similarity calculation on the edited images contained in the editing time series data to obtain a similarity sequence, and perform editing key identification based on the similarity sequence to obtain editing key nodes.
[0081] Step S508: Generate snapshot images of each edit difference node and each edit key node.
[0082] Step S510: Based on the image features of the snapshot image and the image features of the preceding adjacent snapshot image, perform feature difference calculation and difference localization to obtain the editing difference region.
[0083] Step S512: Based on the edit difference region, perform image semantic recognition to obtain the edit difference object and edit difference type, and generate edit prompt words based on the edit difference object and edit difference type.
[0084] Step S514: Perform region of interest detection on the snapshot image to obtain the region of interest.
[0085] Step S516: Determine thumbnail generation parameters based on the region of interest and the region of difference in editing.
[0086] Step S518: Generate thumbnails from snapshot images according to thumbnail generation parameters.
[0087] Step S520: Synchronize thumbnails and editing prompts to the interactive terminal of the intelligent agent.
[0088] It should be noted that any one or more steps in steps S502 to S520 can be combined with any one or more steps in steps S202 to S208 to form a new implementation method according to the needs of implementation and deployment. In addition, any one or more technical features in steps S502 to S520 can be selected and combined with any one or more technical features provided in steps S202 to S208 to form a new implementation method according to the actual deployment needs. Alternatively, any one or more technical features in steps S502 to S520 can be replaced with any one or more technical features provided in steps S202 to S208 to form a new implementation method according to the actual deployment needs. These will not be elaborated on here.
[0089] Furthermore, it should be noted that steps S502 to S520 provided in this embodiment can be executed by the server of the intelligent agent. It should be noted that steps S502 to S520 executed by the server of the intelligent agent and steps S522 to S528 executed by the interactive terminal of the intelligent agent in the following embodiment can cooperate with each other during execution. Therefore, when reading this embodiment, please refer to the corresponding content of steps S522 to S528 provided in the following method embodiment, and when reading the following method embodiment, please refer to the corresponding content of steps S502 to S520 provided in this embodiment.
[0090] One or more embodiments of another image processing method applied to intelligent agents provided in this specification are as follows: Reference Figure 6 The image processing method for intelligent agents provided in this embodiment can be applied to the interactive end of intelligent agents. The method specifically includes steps S602 to S606.
[0091] Step S602: The generated image is edited by the interaction terminal of the intelligent agent and the server.
[0092] In this embodiment, an intelligent agent refers to an entity that can autonomously perform tasks, make decisions, and learn and adjust according to changes in the environment. Specifically, an intelligent agent can be an entity that integrates image generation and image editing capabilities. The intelligent agent can call one or more models, receive natural language instructions or interactive operations input by the user, drive the model to perform corresponding image processing, and manage the state data during the image editing process. Here, the intelligent agent can specifically be an intelligent agent, an intelligent agent application, an intelligent agent system, a large language model (LLM), or the intelligent agent can also be an image generation model.
[0093] The generated image refers to a digital image generated by an intelligent agent through an artificial intelligence algorithm; optionally, the generated image is obtained by the intelligent agent through image generation.
[0094] In practical implementation, the agent's interactive end can cooperate with the server to generate and obtain images. For example, after receiving text prompts submitted by the interactive end, the server can generate an image based on the text prompts using the agent. Furthermore, upon obtaining the generated image, the server can cooperate with the interactive end to edit it. Specifically, when the server receives an editing command for the generated image initiated by the user through the interactive end, it can call an image editing model to perform corresponding image editing processing. During this process, the editing data for each image editing process can be transmitted to the data processing unit to generate editing time-series data. Optionally, the server can call the image editing model only after detecting an editing command for the generated image.
[0095] Here, during the image editing process between the server and the client, the server can synchronize the edited image obtained from the image editing to the client in real time, and the client can also display the synchronized edited image in a synchronized manner.
[0096] The image editing model refers to a model used to edit the generated image, specifically a model that performs multiple rounds of editing on the generated image. The image editing model can be an algorithm model, such as a deep learning-based algorithm model, a model implemented based on a neural network algorithm, or a model implemented based on other algorithms. The input of the image editing model includes the generated image and / or the image editing instructions input by the user, and the output includes image editing data. Here, when the above-mentioned agent is an image generation model, the image editing model and the image generation model that performs image generation can be one model or two models, and this embodiment does not limit this.
[0097] Editing time-series data refers to a set of data related to editing behavior recorded in chronological order or in the order of operations during continuous editing operations on a specific object. Specifically, editing time-series data refers to a set of data related to image editing recorded in chronological order during multiple rounds of editing of the generated image. Editing time-series data may include the edited image, editing archive instructions, editing time and / or editing parameters corresponding to each editing step.
[0098] In the specific execution process, during image processing, in order to ensure data integrity and comprehensiveness, the server reads the editing timing data, that is, reads the editing timing data of the image editing model called to perform image editing on the generated image.
[0099] Step S604: Receive the thumbnail of the snapshot image and status prompt text of the edit status node synchronized by the server.
[0100] In practice, after the server performs multiple rounds of image editing on the generated image based on editing instructions, considering the large number of edited versions and high data redundancy generated by multiple rounds of image editing, in order to avoid low filtering efficiency due to too many invalid edited versions, the server can extract the key editing stages from the editing time sequence data of multiple rounds of image editing. Based on this, the server performs editing state recognition on the editing time sequence data to obtain editing state nodes. On this basis, in order to convert the relevant editing parameters and / or image feature modifications in the editing time sequence data into semantic information, the server further generates status prompt words for the editing state nodes.
[0101] Editing status recognition refers to the process of analyzing editing time-series data to select editing statuses or stages that are significant from the data.
[0102] The edit status node refers to a landmark node that can represent a specific editing stage or editing state; specifically, it refers to a snapshot node in the editing time series data that can be used to represent an important state in the image editing process; for example, the edit status node can be a core object change node in the generated image, or an image style change node in the generated image; here, the edit status node can also be replaced with a snapshot node, and correspondingly, the following description of the edit status node can also be replaced with a snapshot node.
[0103] The status prompt words refer to text used to concisely and accurately describe a specific status, operation content, and / or core information; specifically, they refer to editing prompt words or editing keywords generated based on the editing operation, image changes, and / or editing parameters corresponding to the editing status node. Status prompt words can be used to describe the key or major editing changes in the image of the current editing status node compared to the previous editing status node, so as to inform the user of the core editing content of the current editing status node; here, status prompt words can also be replaced with editing prompt words, editing keywords, or status keywords, and correspondingly, the following descriptions of status prompt words can also be replaced with editing prompt words, editing keywords, or status keywords.
[0104] Furthermore, given the aforementioned determination of edit status nodes, the server can generate snapshot images corresponding to each edit status node. Based on this, in order to accurately characterize the core editing features of the snapshot image compared to the previous snapshot image, and to transform the abstract edit status nodes into intuitive visual elements to support the visualization of the edit status nodes, the server performs image editing detection on the snapshot images based on the editing time-series data, and generates thumbnails of the snapshot images according to the editing detection results. Here, the generated thumbnails of the snapshot images are also known as snapshot thumbnails. Based on this, thumbnails can also be replaced with snapshot thumbnails. Correspondingly, the following descriptions of thumbnails can also be replaced with snapshot thumbnails.
[0105] In this context, a thumbnail refers to a small preview image generated through image processing based on the core editing features of a snapshot image. It is a low-resolution and / or small-size representation of the snapshot image. It should be noted that a thumbnail is not a uniform compression of the entire snapshot image, but rather an image processing method generated based on the cropping and focusing strategies of the snapshot image determined by the editing detection results. Figure 1 On the one hand, it can preserve the main visual content in the image, such as the main object; on the other hand, it can prioritize ensuring that the edited difference area is visible and prominent in the thumbnail. For example, if the user only adds a bird in the lower right corner of the image, even if the area does not occupy the center position in the original composition, the thumbnail will retain the added bird in the thumbnail through cropping and / or layout adjustments, avoiding the loss of key editing information due to regular scaling.
[0106] In specific implementation, when the server obtains the thumbnail and status prompt, the server can synchronize the thumbnail and status prompt to the agent's interactive end. Correspondingly, here, the server receives the thumbnail and status prompt of the snapshot image of the edit status node synchronized by the server.
[0107] Optionally, the edit status node is obtained by performing edit status recognition on the edit time sequence data of the generated image; the thumbnail is generated based on the edit detection results obtained by performing image edit detection on the snapshot image; Step S606: Snapshot synchronization is performed based on the thumbnail and the status prompt.
[0108] In specific implementation, based on the thumbnail and status prompt of the snapshot image of the edit status node synchronized by the receiving server, the snapshot is displayed synchronously here based on the thumbnail and status prompt.
[0109] In the specific execution process, during the snapshot synchronization display on the interactive end, in order to eliminate the comparison error caused by image offset, improve the visualization of differences in edited images, and improve the efficiency of users' version backtracking and comparison operations, the snapshot image corresponding to any thumbnail can be aligned with the preceding adjacent snapshot image, and the aligned snapshot image can be compared and displayed. In one optional implementation of this embodiment, snapshot synchronization display is based on thumbnails and status prompts, including: The image editing interface on the interactive terminal displays a thumbnail of the snapshot image and status prompts; If any thumbnail displayed in the image editing interface is triggered, the snapshot image corresponding to any thumbnail is aligned with the preceding adjacent snapshot image, and the aligned snapshot image is compared and displayed according to the layered display.
[0110] Specifically, after receiving the thumbnail and status prompt of the snapshot image of the edit status node synchronized by the server, the agent's interactive terminal displays the thumbnail and status prompt in the image editing interface of the interactive terminal. Subsequently, if any thumbnail displayed in the image editing interface is triggered, the snapshot image corresponding to the triggering thumbnail can be registered with the preceding adjacent snapshot image at the pixel level and / or feature level, so that the positions of the same background and main object in the two snapshot images are completely overlapped, that is, the target snapshot image corresponding to any triggered thumbnail is image aligned with the reference snapshot image. After that, based on the image alignment, two layers can be created in the editing canvas of the interactive terminal, and the image-aligned target snapshot image and the reference snapshot image can be overlaid and displayed as two independent layers.
[0111] For example, such as Figure 3 The image editing interface shown displays thumbnails 301-1 and status prompts 301-2. If thumbnail 302 is triggered, the snapshot image corresponding to thumbnail 302 is aligned with the preceding adjacent snapshot image; that is, the snapshot image corresponding to thumbnail 302 is aligned with the snapshot image corresponding to thumbnail 303. The image editing interface displays the aligned snapshot images in layers for comparison. Figure 4 As shown, Figure 4The image editing interface shown displays snapshot image 401 corresponding to thumbnail 302 and snapshot image 402 corresponding to thumbnail 303. Snapshot image 401 and snapshot image 402 are pixel-level aligned to ensure that the background textures are completely overlapped. Users can adjust the screen ratio of snapshot image 401 and snapshot image 402 by dragging the "Drag to view differences" control 403. If the progress bar of the "Drag to view differences" control is slid to the left, the display ratio of snapshot image 401 corresponding to thumbnail 302 will gradually increase, while the display ratio of snapshot image 402 on the right will decrease accordingly. If the progress bar of the "Drag to view differences" control is slid to the right, the display ratio of snapshot image 402 will increase. Alternatively, users can also adjust the screen ratio of snapshot image 401 and snapshot image 402 by using control 404.
[0112] It should be added that each optional implementation method and each feasible execution method in steps S602 to S606 provided in this embodiment can be executed independently as needed, or they can be combined and referenced with each other. At the same time, each specific execution step in each optional implementation method or each feasible execution method can also be executed independently or combined as needed. Any feature in each execution step can also be deleted, or any feature in one execution step can be added to another execution step or replace any feature in another execution step. The execution conditions of "if" or "under what circumstances" involved in each step or operation can be directly deleted. This embodiment does not specifically limit the subsequent operations after the execution conditions.
[0113] In summary, the image processing method for intelligent agents provided in this embodiment, during image processing, involves image editing of the generated image through cooperation between the intelligent agent's interactive terminal and the server. Based on this, considering the large number of edited versions and high data redundancy generated by multiple rounds of image editing, and to avoid low filtering efficiency due to an excessive number of invalid edited versions, the server first identifies the edit status nodes in the edit time-series data and generates status prompts for these nodes. Furthermore, it can generate snapshot images corresponding to each edit status node. Subsequently, to accurately characterize the core differences between the snapshot image and the previous snapshot image... The system utilizes heart-editing features and transforms abstract editing state nodes into intuitive visual elements. The server performs image editing detection on snapshot images based on editing time-series data, generates thumbnails of the snapshot images based on the detection results, and synchronizes the thumbnails and status prompts to the agent's interactive end. Subsequently, it receives the thumbnails and status prompts of the snapshot images of the editing state nodes synchronized from the server and displays the snapshots synchronously based on these thumbnails and prompts. This improves the accuracy and intuitiveness of the representation of editing features for each edited image, enabling streamlined and structured management of each edited image, reducing the user's cognitive burden regarding version management of each edited image, and improving the efficiency of users in reviewing and revising edited images.
[0114] The following example uses the image processing method for intelligent agents provided in this embodiment as an example in an image processing scenario, combined with... Figure 5 The image processing method applied to intelligent agents provided in this embodiment will be further described below. Figure 5 An image processing method for intelligent agents, applicable to image processing scenarios, specifically includes the following steps.
[0115] Step S522: Receive the thumbnail of the snapshot image of the edit status node synchronized by the server of the intelligent agent and the edit prompt.
[0116] Optionally, the edit status node includes an edit difference node and / or an edit key node; the edit status node is obtained by performing edit status recognition on the edit time sequence data of the generated image; the thumbnail is generated based on the edit detection results obtained by performing image edit detection on the snapshot image.
[0117] Step S524: Display a thumbnail of the snapshot image and editing prompts in the image editing interface.
[0118] Step S526: If any thumbnail displayed in the image editing interface is triggered, perform image alignment between the snapshot image corresponding to any thumbnail and the preceding adjacent snapshot image.
[0119] Step S528: Compare and display the snapshot images after image alignment according to the layered display method.
[0120] It should be noted that any one or more steps in steps S522 to S528 can be combined with any one or more steps in steps S602 to S606 to form a new implementation method according to the needs of implementation and deployment. In addition, any one or more technical features in steps S522 to S528 can be selected and combined with any one or more technical features provided in steps S602 to S606 to form a new implementation method according to the actual deployment needs. Alternatively, any one or more technical features in steps S522 to S528 can be replaced with any one or more technical features provided in steps S602 to S606 to form a new implementation method according to the actual deployment needs. These will not be elaborated on here.
[0121] This specification provides an embodiment of an image processing device applied to intelligent agents, as follows: In the above embodiments, an image processing method for intelligent agents is provided, and correspondingly, an image processing device for intelligent agents is also provided, which will be described below with reference to the accompanying drawings.
[0122] Reference Figure 7 This illustration shows a schematic diagram of an image processing device embodiment applied to an intelligent agent provided in this embodiment.
[0123] Since the apparatus embodiments correspond to the method embodiments, the descriptions are relatively simple. For relevant parts, please refer to the corresponding descriptions of the method embodiments provided above. The apparatus embodiments described below are merely illustrative.
[0124] This embodiment provides an image processing device for use with an intelligent agent, the device comprising: The data reading module 702 is configured to read the editing timing data of the image editing model called by the intelligent agent's server to perform image editing on the generated image; The node identification module 704 is configured to identify the editing status of the editing time sequence data to obtain editing status nodes, and generate status prompt words for the editing status nodes; The thumbnail generation module 706 is configured to perform image editing detection on the snapshot image based on the editing time series data, and generate a thumbnail of the snapshot image according to the editing detection result; The data synchronization module 708 is configured to synchronize the thumbnail and the status prompt to the interactive terminal of the intelligent agent for snapshot synchronization display.
[0125] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0126] Another embodiment of an image processing device for intelligent agents provided in this specification is as follows: In the above embodiments, another image processing method for intelligent agents is provided, and correspondingly, another image processing apparatus for intelligent agents is also provided, which will be described below with reference to the accompanying drawings.
[0127] Reference Figure 8 This illustration shows a schematic diagram of an image processing device embodiment applied to an intelligent agent provided in this embodiment.
[0128] Since the apparatus embodiments correspond to the method embodiments, the descriptions are relatively simple. For relevant parts, please refer to the corresponding descriptions of the method embodiments provided above. The apparatus embodiments described below are merely illustrative.
[0129] The image editing module 802 is configured to edit the generated image in cooperation with the server through the interactive end of the intelligent agent; The data receiving module 804 is configured to receive a thumbnail and status prompt of the snapshot image of the edit status node synchronized by the server; the edit status node is obtained by performing edit status recognition on the edit time sequence data of the generated image, and the thumbnail is generated based on the edit detection result obtained by performing image edit detection on the snapshot image; The snapshot display module 806 is configured to display snapshots synchronously based on the thumbnail and the status prompt words.
[0130] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0131] This specification provides an example of an image processing device applied to intelligent agents, as follows: Corresponding to the image processing method for an intelligent agent described above, based on the same technical concept, one or more embodiments of this specification also provide an image processing device for an intelligent agent, which is used to execute the image processing method for an intelligent agent provided above. Figure 9 This is a schematic diagram of the structure of an image processing device applied to an intelligent agent, provided for one or more embodiments of this specification.
[0132] This embodiment provides an image processing device for use with intelligent agents, comprising: like Figure 9As shown, device 900 mainly consists of a communication interface 902, a user interface 904, a processor 906, and a data storage 908. These components are interconnected and communicate with each other via a system bus, network, or other connection mechanism 910. Communication interface 902 enables device 900 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, communication interface 902 may include a chipset and antenna for wireless communication with a radio access network or access point. Furthermore, communication interface 902 can be a wired interface such as Ethernet, Token Ring, or USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or wide area wireless interface (e.g., WiMAX or LTE). Of course, communication interface 902 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. Communication interface 902 may also include multiple physical communication interfaces, such as Wi-Fi, Bluetooth, and wide area wireless interfaces. User interface 904 includes receiving user input and providing output to the user. Therefore, user interface 904 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. User interface 904 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, user interface 904 may include software, circuitry, or other forms of logic capable of transmitting data to and receiving data from external user input / output devices. Additionally or alternatively, device 900 may support remote access from other devices via communication interface 902 or another physical interface (not shown). User interface 904 may be configured to receive user input, the position and movement of which may be indicated by indicators or cursors described herein. User interface 904 may also be configured as a display device for rendering or displaying text fragments.
[0133] Processor 906 may include one or more general-purpose processors and / or special-purpose processors. Data storage 908 may include one or more volatile and / or non-volatile storage components, and may be integrated wholly or partially with processor 906. Data storage 908 may include removable and non-removable components.
[0134] Processor 906 is capable of executing program instructions 918 (e.g., compiled or uncompiled program logic and / or machine code) stored in data storage 908 to perform the various functions described herein. Data storage 908 may contain a non-transitory computer-readable medium on which program instructions are stored, which, when executed by device 900, enable device 900 to perform any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Execution of program instructions 918 by processor 906 may result in processor 906 using data 912. For example, program instructions 918 may include an operating system 922 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 900 and one or more application programs 920 (e.g., a browser, social application, or game application). Similarly, data 912 may include operating system data 916 and application data 914. Operating system data 916 is primarily accessible to operating system 922, while application data 914 is primarily accessible to one or more application programs 920. Application data 914 may reside in a file system visible or hidden from the user of device 900. Application 920 can communicate with operating system 922 through one or more application programming interfaces (APIs). These APIs facilitate application 920 reading and / or writing application data 914, transmitting or receiving information via communication interface 902, and receiving or displaying information on user interface 904. In some terms, application 920 may be simply referred to as "app". Furthermore, application 920 can be downloaded to device 900 through one or more online application stores or app markets. However, applications can also be installed on device 900 in other ways, such as through a web browser or a physical interface on device 900 (e.g., a USB port).
[0135] In one specific embodiment, the image processing device applied to an intelligent agent includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the image processing device applied to the intelligent agent, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following: Read the editing timing data of the image editing model called by the intelligent agent's server to perform image editing on the generated image; Editing status nodes are obtained by performing editing status identification on the editing time sequence data, and status prompt words for the editing status nodes are generated; Image editing detection is performed on the snapshot image based on the editing time-series data, and a thumbnail of the snapshot image is generated based on the editing detection results; The thumbnail and status prompt are synchronized to the interactive terminal of the intelligent agent to perform snapshot synchronization display.
[0136] Another embodiment of an image processing device for intelligent agents provided in this specification is as follows: Corresponding to the image processing method for intelligent agents described above, based on the same technical concept, one or more embodiments of this specification also provide another image processing apparatus for intelligent agents, which is used to execute the other image processing method for intelligent agents provided above. Figure 10 This is a schematic diagram of another image processing device applied to an intelligent agent, provided for one or more embodiments of this specification.
[0137] This embodiment provides an image processing device for use with intelligent agents, comprising: like Figure 10As shown, device 1000 mainly consists of a communication interface 1002, a user interface 1004, a processor 1006, and a data storage 1008. These components are interconnected and communicate with each other via a system bus, network, or other connection mechanism 1010. The communication interface 1002 enables device 1000 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface 1002 may include a chipset and antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface 1002 can be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface 1002 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface 1002 may also include multiple physical communication interfaces, such as Wi-Fi, Bluetooth, and wide-area wireless interfaces. The user interface 1004 includes receiving user input and providing output to the user. Therefore, the user interface 1004 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. The user interface 1004 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, the user interface 1004 may include software, circuitry, or other forms of logic capable of transmitting data to and receiving data from external user input / output devices. Additionally or alternatively, the device 1000 may support remote access from other devices via communication interface 1002 or another physical interface (not shown). The user interface 1004 may be configured to receive user input, the position and movement of which may be indicated by indicators or cursors described herein. The user interface 1004 may also be configured as a display device for rendering or displaying text fragments.
[0138] Processor 1006 may include one or more general-purpose processors and / or dedicated processors. Data storage 1008 may include one or more volatile and / or non-volatile storage components, and may be integrated wholly or partially with processor 1006. Data storage 1008 may include removable and non-removable components.
[0139] Processor 1006 is capable of executing program instructions 1018 (e.g., compiled or uncompiled program logic and / or machine code) stored in data storage 1008 to perform the various functions described herein. Data storage 1008 may contain a non-transitory computer-readable medium on which program instructions are stored, which, when executed by device 1000, enable device 1000 to perform any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Execution of program instructions 1018 by processor 1006 may result in processor 1006 using data 1012. For example, program instructions 1018 may include an operating system 1022 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 1000 and one or more application programs 1020 (e.g., a browser, social application, or game application). Similarly, data 1012 may include operating system data 1016 and application data 1014. Operating system data 1016 is primarily accessible to operating system 1022, while application data 1014 is primarily accessible to one or more application programs 1020. Application data 1014 may reside in a file system visible or hidden by the user of device 1000. Application 1020 may communicate with operating system 1022 via one or more application programming interfaces (APIs). These APIs facilitate application 1020 reading and / or writing application data 1014, transmitting or receiving information via communication interface 1002, receiving or displaying information on user interface 1004, etc. In some terms, application 1020 may be simply referred to as an "app". Furthermore, application 1020 may be downloaded to device 1000 through one or more online app stores or app markets. However, applications may also be installed on device 1000 in other ways, such as through a web browser or a physical interface on device 1000 (e.g., a USB port).
[0140] In one specific embodiment, the image processing device applied to an intelligent agent includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the image processing device applied to the intelligent agent, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following: The generated image is edited by the interaction between the intelligent agent's interactive terminal and the server. The system receives thumbnails and status prompts of snapshot images from the edit status nodes synchronized by the server. The edit status nodes are obtained by performing edit status recognition on the edit time sequence data of the generated image, and the thumbnails are generated based on the edit detection results obtained by performing image edit detection on the snapshot images. A snapshot is displayed synchronously based on the thumbnail and the status prompt.
[0141] This specification provides an embodiment of a computer-readable storage medium as follows: Corresponding to the image processing method for intelligent agents described above, based on the same technical concept, one or more embodiments of this specification also provide a computer-readable storage medium.
[0142] The computer-readable storage medium provided in this embodiment is used to store computer-executable instructions, which, when executed, implement the following process: Read the editing timing data of the image editing model called by the intelligent agent's server to perform image editing on the generated image; Editing status nodes are obtained by performing editing status identification on the editing time sequence data, and status prompt words for the editing status nodes are generated; Image editing detection is performed on the snapshot image based on the editing time-series data, and a thumbnail of the snapshot image is generated based on the editing detection results; The thumbnail and status prompt are synchronized to the interactive terminal of the intelligent agent to perform snapshot synchronization display.
[0143] It should be noted that the embodiments of a computer-readable storage medium described in this specification and the embodiments of an image processing method applied to an intelligent agent described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.
[0144] Another embodiment of a computer-readable storage medium provided in this specification is as follows: In response to another image processing method for intelligent agents described above, based on the same technical concept, one or more embodiments of this specification also provide another computer-readable storage medium.
[0145] The computer-readable storage medium provided in this embodiment is used to store computer-executable instructions, which, when executed, implement the following process: The generated image is edited by the interaction between the intelligent agent's interactive terminal and the server. The system receives thumbnails and status prompts of snapshot images from the edit status nodes synchronized by the server. The edit status nodes are obtained by performing edit status recognition on the edit time sequence data of the generated image, and the thumbnails are generated based on the edit detection results obtained by performing image edit detection on the snapshot images. A snapshot is displayed synchronously based on the thumbnail and the status prompt.
[0146] It should be noted that the embodiments of another computer-readable storage medium described in this specification and the embodiments of another image processing method applied to intelligent agents described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.
[0147] This specification provides an example of a computer program product as follows: Corresponding to the image processing method applied to intelligent agents described above, based on the same technical concept, one or more embodiments of this specification also provide a computer program product.
[0148] A computer program product includes a computer program / instructions that, when executed by a processor, perform the following steps: Read the editing timing data of the image editing model called by the intelligent agent's server to perform image editing on the generated image; Editing status nodes are obtained by performing editing status identification on the editing time sequence data, and status prompt words for the editing status nodes are generated; Image editing detection is performed on the snapshot image based on the editing time-series data, and a thumbnail of the snapshot image is generated based on the editing detection results; The thumbnail and status prompt are synchronized to the interactive terminal of the intelligent agent to perform snapshot synchronization display.
[0149] It should be noted that the embodiments of a computer program product described in this specification and the embodiments of an image processing method applied to an intelligent agent described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.
[0150] Another example of a computer program product provided in this specification is as follows: Corresponding to the image processing method applied to intelligent agents described above, based on the same technical concept, one or more embodiments of this specification also provide another computer program product.
[0151] A computer program product includes a computer program / instructions that, when executed by a processor, perform the following steps: The generated image is edited by the interaction between the intelligent agent's interactive terminal and the server. The system receives thumbnails and status prompts of snapshot images from the edit status nodes synchronized by the server. The edit status nodes are obtained by performing edit status recognition on the edit time sequence data of the generated image, and the thumbnails are generated based on the edit detection results obtained by performing image edit detection on the snapshot images. A snapshot is displayed synchronously based on the thumbnail and the status prompt.
[0152] It should be noted that the embodiments of another computer program product described in this specification and the embodiments of another image processing method applied to intelligent agents described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.
[0153] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments. For example, the device embodiment, equipment embodiment and computer-readable storage medium embodiment are all similar to the method embodiment, so the description is relatively simple. When reading the relevant content of the device embodiment, equipment embodiment and computer-readable storage medium embodiment, please refer to the description of the method embodiment.
[0154] While one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is merely one possible execution order among many steps, and does not represent the only execution order. Therefore, when the claims involve method steps, any changes or adjustments to the order of such steps, or the parallelism between steps, are also within the scope of protection of the claims. This specification uses specific terms to describe embodiments of this specification. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic related to at least one embodiment of this specification. Therefore, it should be emphasized and noted that "an embodiment," "one embodiment," or "an alternative embodiment" mentioned twice or more in different locations in this specification do not necessarily refer to the same embodiment. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0155] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0156] In the 1930s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many improvements to the methodology today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that an improvement to the methodology cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0157] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0158] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0159] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, when implementing the embodiments of this specification, the functions of each unit can be implemented in one or more software and / or hardware.
[0160] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-readable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0161] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0162] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0163] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0164] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0165] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0166] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer-readable storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0167] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising at least one…" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0168] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0169] The above description is merely an embodiment of this document and is not intended to limit the scope of this document. Various modifications and variations can be made to this document by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this document should be included within the scope of the claims of this document.
Claims
1. An image processing method applied to intelligent agents, comprising: Read the editing timing data of the image editing model called by the intelligent agent's server to perform image editing on the generated image; Editing status nodes are obtained by performing editing status identification on the editing time sequence data, and status prompt words for the editing status nodes are generated; Image editing detection is performed on the snapshot image based on the editing time-series data, and a thumbnail of the snapshot image is generated based on the editing detection results; The thumbnail and status prompt are synchronized to the interactive terminal of the intelligent agent to perform snapshot synchronization display.
2. The image processing method applied to an intelligent agent according to claim 1, wherein the step of obtaining edit state nodes by performing edit state recognition on the edit time-series data includes: Edit difference detection is performed on the edited images contained in the edit time series data to obtain edit difference nodes; The similarity sequence is obtained by performing similarity calculation on the edited images contained in the editing time series data, and the editing key nodes are obtained by performing editing key identification based on the similarity sequence.
3. The image processing method applied to an intelligent agent according to claim 1, wherein the step of obtaining edit state nodes by performing edit state recognition on the edit time-series data includes: Visual and semantic features are extracted and merged from each edited image contained in the editing time series data to obtain visual and semantic features; Calculate the feature difference value between the visual semantic features of the index image and the visual semantic features of the main index image in each edited image, and determine the edit state node based on the obtained feature difference value.
4. The image processing method applied to an intelligent agent according to claim 3, wherein determining the edit state node based on the obtained feature difference value includes: The significance value of the difference is obtained by calculating the significance of the difference between the edited images based on the feature difference values; The edit node corresponding to the edited image with a significant difference value greater than the adaptive threshold is determined as the edited state node; the adaptive threshold is determined based on the numerical distribution index of the significant difference values of each edited image.
5. The image processing method applied to an intelligent agent according to claim 1, wherein the step of obtaining edit state nodes by performing edit state recognition on the edit time-series data includes: Feature similarity sequences are obtained by performing feature similarity calculations on adjacent edited images contained in the edit time series data; The similarity subsequences with a specific editing pattern are detected in the feature similarity sequence, and the editing pattern of the specific editing pattern is verified on the similarity subsequences. The editing state node is determined in the similarity subsequences that pass the verification.
6. The image processing method applied to an intelligent agent according to claim 1, after the step of reading the editing sequence data of the image editing model called by the server of the intelligent agent to edit the generated image is executed, and before the operation of generating the status prompt word of the editing status node is executed, it further includes: The editing status node is determined based on the editing archive instructions contained in the editing timing data, and / or the editing status node is determined based on the editing time and editing status cycle contained in the editing timing data.
7. The image processing method applied to an intelligent agent according to any one of claims 2 to 6, wherein generating the status prompt word of the edit state node includes: Based on the image features of the snapshot image and the image features of the preceding adjacent snapshot image, feature difference calculation and difference localization are performed to obtain the editing difference region; Based on the edit difference region, image semantic recognition is performed to obtain the edit difference object and / or edit difference type, and edit prompt words are generated based on the edit difference object and / or the edit difference type.
8. The image processing method applied to an intelligent agent according to any one of claims 2 to 6, wherein generating the status prompt word of the edit state node includes: The editing parameters of each edited image in the editing time segment corresponding to the editing state node are read, and the read editing parameters are fused to obtain editing parameter prompts.
9. The image processing method applied to an intelligent agent according to claim 1, wherein performing image edit detection on the snapshot image based on the edit timing data and generating a thumbnail of the snapshot image according to the edit detection result includes: Regions of interest are obtained by performing region of interest detection on the snapshot image, and edit difference recognition is performed on the snapshot image to obtain edit difference regions; Based on the region of interest and the region of difference in editing, the thumbnail generation parameters are determined, and the snapshot image is processed according to the thumbnail generation parameters to generate the thumbnail.
10. The image processing method applied to an intelligent agent according to claim 9, wherein performing image edit detection on the snapshot image based on the edit timing data and generating a thumbnail of the snapshot image according to the edit detection result includes: The snapshot image is subjected to edit object identification and / or edit region detection, and the obtained edit object and / or edit region is segmented to obtain target image blocks; The thumbnail is obtained by explicitly marking the target image block according to the explicit marking method corresponding to the editing type of the editing object and / or editing area.
11. The image processing method applied to an intelligent agent according to claim 1, wherein the snapshot synchronization display is implemented in the following manner: The thumbnail and the status prompt are displayed in the image editing interface of the interactive terminal; If any thumbnail displayed in the image editing interface is triggered, the snapshot image corresponding to the thumbnail is aligned with the preceding adjacent snapshot image, and the aligned snapshot image is compared and displayed according to the layered display.
12. The image processing method applied to an intelligent agent according to claim 1, wherein the generated image is obtained by the intelligent agent generating the image; the invocation of the image editing model is performed after detecting an editing instruction for the generated image; The image processing method applied to the intelligent agent is applied to the image editing model or snapshot generation model; The snapshot generation model performs snapshot generation processing in response to the call of the image editing model.
13. An image processing method applied to an intelligent agent, comprising: The generated image is edited by the interaction between the intelligent agent's interface and the server. Receive thumbnails and status prompts of snapshot images of edit status nodes synchronized by the server; The edit status node is obtained by performing edit status recognition on the edit time sequence data of the generated image; the thumbnail is generated based on the edit detection results obtained by performing image edit detection on the snapshot image; A snapshot is displayed synchronously based on the thumbnail and the status prompt.
14. The image processing method applied to an intelligent agent according to claim 13, wherein the snapshot synchronization display based on the thumbnail and the status prompt word includes: The image editing interface on the interactive terminal displays a thumbnail of the snapshot image and the status prompt. If any thumbnail displayed in the image editing interface is triggered, the snapshot image corresponding to the thumbnail is aligned with the preceding adjacent snapshot image, and the aligned snapshot image is compared and displayed according to the layered display.
15. An image processing apparatus for use with intelligent agents, comprising: The data reading module is configured to read the editing time sequence data of the image editing model called by the intelligent agent's server to perform image editing on the generated image; The node identification module is configured to identify the editing status of the editing time sequence data to obtain editing status nodes, and generate status prompt words for the editing status nodes; The thumbnail generation module is configured to perform image editing detection on the snapshot image based on the editing time-series data, and generate a thumbnail of the snapshot image based on the editing detection result; The data synchronization module is configured to synchronize the thumbnail and the status prompt to the interactive terminal of the intelligent agent for snapshot synchronization display.
16. An image processing apparatus for use with intelligent agents, comprising: The image editing module is configured to edit the generated image in cooperation with the server through the interactive end of the intelligent agent. The data receiving module is configured to receive thumbnails and status prompts of snapshot images of the edit status nodes synchronized by the server. The edit status node is obtained by performing edit status recognition on the edit time sequence data of the generated image, and the thumbnail is generated based on the edit detection result obtained by performing image edit detection on the snapshot image; The snapshot display module is configured to display snapshots synchronously based on the thumbnails and the status prompts.
17. An image processing device for use with intelligent agents, comprising: processor; And, a memory configured to store computer-executable instructions, which, when executed, cause the processor to: Read the editing timing data of the image editing model called by the intelligent agent's server to perform image editing on the generated image; Editing status nodes are obtained by performing editing status identification on the editing time sequence data, and status prompt words for the editing status nodes are generated; Image editing detection is performed on the snapshot image based on the editing time-series data, and a thumbnail of the snapshot image is generated based on the editing detection results; The thumbnail and status prompt are synchronized to the interactive terminal of the intelligent agent to perform snapshot synchronization display.
18. An image processing device for use with intelligent agents, comprising: processor; And, a memory configured to store computer-executable instructions, which, when executed, cause the processor to: The generated image is edited by the interaction between the intelligent agent's interface and the server. Receive thumbnails and status prompts of snapshot images of edit status nodes synchronized by the server; The edit status node is obtained by performing edit status recognition on the edit time sequence data of the generated image, and the thumbnail is generated based on the edit detection result obtained by performing image edit detection on the snapshot image; A snapshot is displayed synchronously based on the thumbnail and the status prompt.
19. A computer-readable storage medium for storing computer-executable instructions that, when executed, implement the steps of the method of claim 1 or claim 13.