Text marking method and device, computer equipment, storage medium and program product
By providing text tag pages and multi-modal graphic and text models in social software, and automatically identifying and rendering the image subject objects, the intelligent and diversified needs of graphic and text editing in the existing technology are solved, and the user experience and marking effect are improved.
Patent Information
- Application Number
- CN202510384966.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-11
AI Technical Summary
The existing social software graphic editing functions cannot meet users' intelligent and diversified needs for graphic information, and they need to manually switch the marking function for text marking.
Provide a text marking method, which automatically obtains and renders text marking content by displaying text marking pages, including text input, generation and selection controls, and uses multi-modal graphic and text model and image feature matching technology to intelligently identify the subject objects in the image and prioritize them, so as to realize customization, one-click and intelligent marking.
It realizes the intelligent and diversified marking of graphic and text information, improves user experience, meets personalized needs, and optimizes the rationality of marking effects.
Smart Images

Figure CN120297232A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of application development technologies, and particularly to a text marking method, apparatus, computer device, storage medium, and program product. Background Art
[0002] With the continuous development of Internet applications, more and more people post graphic and text information on social software to share their lives. Therefore, the ways of performing personalized processing on graphic and text information have become diverse.
[0003] Currently, the graphic and text editing functions provided by conventional social software require switching to the marking function after invoking the text panel in the editing scenario for publishing, and manually marking text on the picture. This method cannot meet the intelligent and diverse needs of users for publishing graphic and text information. Summary of the Invention
[0004] Based on this, it is necessary to provide a text marking method, apparatus, computer device, storage medium, and program product that can intelligently perform text marking on images for the above technical problems.
[0005] In a first aspect, this application provides a text marking method, including:
[0006] Display a text marking page; the text marking page displays an initial image to be marked, and the text marking page includes text marking controls; wherein, the text marking controls include at least one of a text input control, a text generation control, and a text selection control;
[0007] In response to a trigger operation on the text marking control, obtain text marking content;
[0008] Generate a target image corresponding to the initial image, where the target image includes the text marking content.
[0009] In one embodiment, in response to a trigger operation on the text marking control, obtaining text marking content includes:
[0010] In response to a trigger operation on the text input control, obtain the text marking content input through the text input control;
[0011] Render the text marking content to obtain the initial image marked based on the text marking content.
[0012] In one embodiment, in response to a trigger operation on the text marking control, obtaining text marking content includes:
[0013] In response to a triggering operation on a text marking control, perform category detection on the initial image to be marked to obtain the categories of the main objects included in each candidate detection box in the initial image; wherein, each candidate detection box contains one main object.
[0014] Determine the priority information of each candidate detection box according to the categories of the main objects included in each candidate detection box.
[0015] Determine the target detection box from each candidate detection box according to the priority information.
[0016] Render the text marking content corresponding to the target detection box to obtain the initial image marked based on the text marking content.
[0017] In one embodiment, when the text marking control is a text generation control, determining the priority information of each candidate detection box according to the categories of the main objects included in each candidate detection box includes:
[0018] Obtain the text marking content corresponding to each candidate detection box in the initial image.
[0019] Determine the priority information of each candidate detection box according to the text marking content corresponding to each candidate detection box and the categories of the main objects included in each candidate detection box.
[0020] In one embodiment, obtaining the text marking content corresponding to each candidate detection box in the initial image includes:
[0021] Extract the image features of each candidate detection box and the text features of the template marking content in the initial image through a multimodal image-text model.
[0022] Match the image features and text features one by one to obtain an image-text matching result.
[0023] Determine the text marking content corresponding to each candidate detection box from the template marking content according to the image-text matching result.
[0024] In one embodiment, determining the priority information of each candidate detection box according to the text marking content corresponding to each candidate detection box and the categories of the main objects included in each candidate detection box includes:
[0025] Determine the category of the initial image according to the text marking content corresponding to each candidate detection box and the categories of the main objects included in each candidate detection box.
[0026] Match the category of the initial image with the categories of the main objects included in each candidate detection box one by one to obtain a category matching result.
[0027] Determine the priority information of each candidate detection box according to the category matching result.
[0028] In one embodiment, when the text marking control is a text selection control, the priority information of each candidate detection box is determined according to the category of the main object included in each candidate detection box, including:
[0029] The confidence information corresponding to each candidate detection box is determined according to the category of the main object included in each candidate detection box;
[0030] The priority information of each candidate detection box is determined according to the confidence information and position information corresponding to each candidate detection box.
[0031] In one embodiment, the priority information of each candidate detection box is determined according to the confidence information and position information corresponding to each candidate detection box, including:
[0032] For each candidate detection box, the confidence information and position information corresponding to the candidate detection box are weighted according to the confidence weight and position weight to obtain the priority information of the candidate detection box.
[0033] In one embodiment, rendering the text marking content corresponding to the target detection box to obtain an initial image marked based on the text marking content includes:
[0034] Generating an updated text selection control according to the text marking content corresponding to the target detection box;
[0035] In response to a trigger operation on the updated text selection control, determining the target text marking content from the text marking content corresponding to the target detection box;
[0036] Rendering the target text marking content to obtain an initial image marked based on the text marking content.
[0037] In one embodiment, before displaying the text marking page, the method further includes:
[0038] Displaying an image editing page; the image editing page displays an initial image, and the image editing page includes a text editing control;
[0039] Correspondingly, displaying the text marking page includes:
[0040] In response to a trigger operation on the text editing control, displaying the text marking page.
[0041] In one embodiment, before generating the target image corresponding to the initial image in response to a trigger operation on the marking completion control, the method further includes:
[0042] In response to a triggering operation on an update control for text-marked content, display a preview page of the updated effect; the preview page of the updated effect displays an updated initial image, and the update control includes at least one of an edit update control, a mirror update control, and a delete update control.
[0043] In a second aspect, the present application also provides a text marking device, including:
[0044] A marking page display module, configured to display a text marking page; the text marking page displays an initial image to be marked, and the text marking page includes text marking controls; wherein, the text marking controls include at least one of a text input control, a text generation control, and a text selection control;
[0045] An acquisition module, configured to acquire text marking content in response to a triggering operation on the text marking control;
[0046] An image generation module, configured to generate a target image corresponding to the initial image, and the target image includes the text marking content.
[0047] In one embodiment, the preview page display module includes:
[0048] A category detection unit, configured to perform category detection on the initial image to be marked in response to a triggering operation on the text marking control, and obtain the category of the main object included in each candidate detection box in the initial image; wherein, each candidate detection box includes one main object;
[0049] A priority determination unit, configured to determine the priority information of each candidate detection box according to the category of the main object included in each candidate detection box;
[0050] A detection box determination unit, configured to determine a target detection box from each candidate detection box according to the priority information;
[0051] A marking rendering unit, configured to render the text marking content corresponding to the target detection box to obtain the initial image marked based on the text marking content.
[0052] In a third aspect, the present application also provides a computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the method in the first aspect are implemented.
[0053] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method in the first aspect are implemented.
[0054] In a fifth aspect, the present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the method in the first aspect described above.
[0055] The above text marking method, device, computer device, storage medium and program product display a text marking page; the text marking page displays an initial image to be marked, and the text marking page includes text marking controls; among them, the text marking controls include at least one of a text input control, a text generation control and a text selection control; in response to a triggering operation on the text marking control, text marking content is obtained; a target image corresponding to the initial image is generated, and the target image includes the text marking content. The present application provides a visual interaction process for text marking. By triggering different controls in the text marking page, the steps of performing functions of custom marking, one-key marking or intelligent marking are executed. Compared with the prior art, the present embodiment provides diversified and intelligent marking methods, which can bring a better marking experience to users and ensure the rationality of the marking effect at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0057] Figure 1 It is an application environment diagram of the text marking method in one embodiment;
[0058] Figure 2 It is a flow diagram of the text marking method in one embodiment;
[0059] Figure 3 It is a page diagram of the text marking page in one embodiment;
[0060] Figure 4 It is a page diagram of the effect preview page in one embodiment;
[0061] Figure 5 It is a flow diagram of the text marking method in another embodiment;
[0062] Figure 6 It is a flow diagram of determining priority information in one embodiment;
[0063] Figure 7 It is a flow diagram of obtaining text marking content in one embodiment;
[0064] Figure 8Schematic flowchart of determining priority information in another embodiment;
[0065] Figure 9 Schematic flowchart of determining priority information in yet another embodiment;
[0066] Figure 10 Schematic flowchart of content rendering in one embodiment;
[0067] Figure 11 Schematic diagram of an image editing page in one embodiment;
[0068] Figure 12 Schematic diagram of an effect preview page in another embodiment;
[0069] Figure 13 Block diagram of a text marking device in one embodiment;
[0070] Figure 14 Block diagram of a text marking device in another embodiment;
[0071] Figure 15 Internal structure diagram of a computer device in one embodiment. Detailed implementation manners
[0072] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0073] The text marking method provided by the embodiments of the present application can be applied to an application environment as shown in Figure 1 . Among them, the terminal device 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed in the cloud or on other network servers. The terminal device 102 can independently execute the text marking method, or can perform data interaction with the server 104, instruct the server 104 to execute the text marking method, and obtain the execution result fed back by the server 104.
[0074] Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server 104 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0075] In an exemplary embodiment, as Figure 2 shown, a text marking method is provided, including:
[0076] S201, display a text marking page.
[0077] As Figure 3 shown, the text marking page displays an initial image to be marked. The initial image can be uploaded by the user to the terminal device, such as selected from the album of the terminal device, or can be taken in real time by the user through the terminal device, such as taken through the camera function.
[0078] The text marking page includes text marking controls; among them, the text marking controls include at least one of a text input control, a text generation control, and a text selection control. The text marking content generally represents a sentence or a group of words to share the user's life.
[0079] Specifically, the text input control corresponds to the custom marking function and can provide a window for the user to input the text marking content, such as a text input box; the text generation control corresponds to the one-key marking function and is a control represented as "one-key marking", which can provide a one-key text marking method for the user without any additional operations by the user; the text selection control corresponds to the intelligent marking function and can be a control represented as "beautiful green tourism day", "lazy posture", or "essential for outdoor activities", which can provide selectable template marking content that is as close as possible to the user's intention.
[0080] S202, in response to a triggering operation on the text marking control, obtain the text marking content.
[0081] Optionally, when obtaining the text marking content in response to a triggering operation on the text marking control, the text marking content can be obtained first, and then the initial image marked based on the text marking content can be obtained.
[0082] Optionally, in the case of obtaining the initial image marked based on the text marking content, a preview page of the effect can be displayed, and the preview page of the effect displays the initial image marked based on the text marking content.
[0083] As Figure 4 shown, the preview page of the effect displays the initial image marked based on the text marking content. Optionally, the preview page of the effect may include a marking completion control. The marking completion control may be a control represented as "OK". Specifically, in response to the triggering operation on the text marking control, the terminal device obtains the corresponding text marking content and renders the text marking content, and then the initial image marked based on the text marking content can be obtained.
[0084] Optionally, the preview page of the effect further includes a marking incomplete control, which may be a control represented as "Cancel".
[0085] Optionally, the preview page of the effect further includes a marking prompt message, which may be represented as "Click on the marked content to edit or delete".
[0086] S203. Generate a target image corresponding to the initial image, where the target image includes the text marking content.
[0087] Optionally, when generating the target image corresponding to the initial image, in response to the triggering operation on the marking completion control, the target image corresponding to the initial image can be generated accordingly.
[0088] When the user clicks the marking completion control, correspondingly, in response to the triggering operation on the marking completion control, the terminal device uses the initial image marked based on the text marking content in the current preview page of the effect as the target image, generates the corresponding image data, and jumps to the next-level page or the previous-level image editing page.
[0089] Optionally, when the user clicks the marking incomplete control, correspondingly, in response to the triggering operation on the marking incomplete control, the terminal device returns to the previous-level text marking page to meet the user's need to perform text marking again.
[0090] In this embodiment, a visual interaction process of a text marking method is provided. By triggering different controls in the text marking page, the steps of performing the functions of custom marking, one-key marking, or intelligent marking are executed. Compared with the prior art, this embodiment provides diversified and intelligent marking methods, which can bring a better marking experience to users and ensure the rationality of the marking effect at the same time.
[0091] In an exemplary embodiment, when the user selects the custom marking function, S202 described above includes:
[0092] In response to a trigger operation on the text input control, obtain the text markup content input via the text input control; render the text markup content to obtain an initial image marked based on the text markup content.
[0093] In response to a trigger operation on the text input control, the terminal device pops up a corresponding text input box on top of the initial image displayed in the text markup page for the user to input the desired text markup content. The text input box includes various controls with different functions, such as adjusting font, color, case, and markup style, etc.
[0094] Furthermore, the terminal device obtains the text markup content input via the text input control, renders the text markup content, and displays it in the initial image on the effect preview page.
[0095] In this embodiment, when the user selects the custom markup function, the terminal device can directly obtain the text markup content input by the user through the text input control and complete the markup process.
[0096] In an exemplary embodiment, as Figure 5 shown, a text markup method is provided, including:
[0097] S501, in response to a trigger operation on the text markup control, perform category detection on the initial image to be marked to obtain the categories of the main objects included in each candidate detection box in the initial image.
[0098] Detect the initial image through the basic detection model to obtain the positions and categories of the main objects in the initial image. During the detection process, generate each candidate detection box corresponding to the initial image. Each candidate detection box contains one main object, and there may be overlapping parts between the candidate detection boxes. Among them, the main object can be an object such as a person, animal, vehicle, and building in the real or virtual scene.
[0099] Perform category detection on each candidate detection box to obtain the category detection results of the main objects included therein. Generally, the category detection results include multiple possible categories and the confidence information corresponding to each category. The category with the highest confidence indicated by the confidence information is used as the category of the main object, and the confidence information corresponding to this category is used as the confidence information corresponding to the candidate detection box.
[0100] Optionally, use the position of the candidate detection box as the position of the main object included in the candidate detection box.
[0101] S502, according to the categories of the main objects included in each candidate detection box, determine the priority information of each candidate detection box.
[0102] It is possible to identify the intention of the user's posted image according to the categories of the subject objects included in each candidate detection box, so as to determine the priority information of each candidate detection box, so as to ensure as much as possible that the subject objects included in the candidate detection boxes with higher priorities are the subject objects emphasized by the user's intention.
[0103] Optionally, the categories of the subject objects included in each candidate detection box in the initial image include people, vehicles, and buildings. Different categories correspond to different priorities. For example, the priority of people is higher than that of vehicles and buildings. Therefore, it is determined that the priority of the candidate detection box containing people is higher than that of the candidate detection box containing vehicles or buildings, and the priorities of each detection box are sorted to obtain the corresponding priority information.
[0104] Optionally, according to the categories of the subject objects included in each candidate detection box and the sizes of each candidate detection box, the priority information of each candidate detection box is determined. It can be understood that the larger the candidate detection box, the larger the subject object included in the candidate detection box, and thus the greater the probability that the subject object is the subject object emphasized by the user's intention. Exemplarily, corresponding weights are assigned to different categories, and combined with the size of the candidate detection box, the corresponding priority information is calculated.
[0105] Optionally, when detecting the initial image, the positions and categories of the subject objects included in each candidate detection box are determined, and the priority information of each candidate detection box can be determined by combining the position and the category. It can be understood that the closer the position of the subject object included in the candidate detection box is to the center of the image, the greater the probability that the subject object is the subject object emphasized by the user's intention. Exemplarily, corresponding weights are assigned to different categories, and combined with the position of the candidate detection box, the corresponding priority information is calculated.
[0106] Optionally, the terminal device responds to the image category selected by the user, such as people, vehicles, or buildings, etc., and then determines that the priority of the candidate detection box that matches the category is higher, and obtains the priority information of each candidate detection box.
[0107] S503. Determine the target detection box from each candidate detection box according to the priority information.
[0108] According to the priority information, arrange each candidate detection box, and select a certain number of candidate detection boxes with higher priorities from them as the target detection box.
[0109] Exemplarily, the candidate detection boxes whose priorities indicated by the priority information are ranked in the top three are used as the target detection boxes. The number of target detection boxes selected in this embodiment is not limited.
[0110] S504. Render the text marking content corresponding to the target detection box to obtain the initial image marked based on the text marking content.
[0111] Among them, the text marking content corresponding to the target detection box is used to characterize the features of the main object included in the target detection box, and is generally expressed as a sentence or a group of words to share the user's life.
[0112] Exemplarily, when the main object included in the target detection box is a person, the text marking content corresponding to the target detection box can be expressed as "lazy posture" or "when going out to play on weekends", and when the main object included in the target detection box is a scenery, the text marking content corresponding to the target detection box can be expressed as "beautiful green tourism day", etc.
[0113] In this embodiment, the text marking content corresponding to the target detection box can be determined in various ways. Specifically, the function of text marking for the initial image has at least three implementation methods, namely custom marking, one-key marking, and intelligent marking.
[0114] When the user selects the one-key marking function, in the text marking page, click the text generation control, that is, the control represented as "one-key marking". Correspondingly, the terminal device responds to the trigger operation on the text generation control, renders the text marking content corresponding to the target detection box, and obtains the marked initial image. Specifically, before or after the terminal device senses the trigger operation on the text generation control, it can execute the above steps S501 - S503, and select the text marking content corresponding to the target detection box from the preset template marking content.
[0115] When the user selects the intelligent marking function, in the text marking page, click the text selection control corresponding to the desired text marking content, such as the control represented as "beautiful green tourism day", "lazy posture", or "essential for outdoor activities". Correspondingly, the terminal device responds to the trigger operation on the text selection control, renders the text marking content selected by the user, and obtains the marked initial image. Specifically, the terminal device can first display the initial text selection control to the user, and then execute the above steps S501 - S503 after responding to the trigger operation of the user on the text selection control, generate the text selection control corresponding to the text marking content of the target detection box, and complete the update of the text selection control. The terminal device can also execute the above steps S501 - S503 after obtaining the initial image to be marked, select the text marking content corresponding to the target detection box from the preset template marking content, and then generate the text selection control corresponding to the text marking content for the user to select.
[0116] In the above text marking method, category detection is performed on the initial image to be marked to obtain the categories of the main objects included in each candidate detection box in the initial image; wherein, each candidate detection box includes one main object; according to the categories of the main objects included in each candidate detection box, determine the priority information of each candidate detection box; according to the priority information, determine the target detection box from each candidate detection box; render the text marking content corresponding to the target detection box to obtain the marked initial image. In this embodiment, by identifying the categories of different main objects in the initial image, the corresponding priority information is determined, so as to achieve the purpose of perceiving the intention of the user to publish text and image information, determine the text marking content that better conforms to the user's intention for rendering, and realize the intelligent text marking of the initial image without manual marking, which can not only improve the personalized marking effect, but also meet different marking requirements to a large extent.
[0117] In an exemplary embodiment, as Figure 6 shown, when the user selects the one-key marking function, S502 described above includes:
[0118] S601, obtain the text marking content corresponding to each candidate detection box in the initial image.
[0119] Through an artificial intelligence / machine learning model, generate the text marking content corresponding to each candidate detection box. Among them, the artificial intelligence / machine learning model can include multiple sub-models to enable the artificial intelligence / machine learning model to process image content and text content simultaneously.
[0120] Specifically, the initial image can be input into a basic detection model to determine each candidate detection box in the initial image, and then determine the category of the main object included in each candidate detection box, perform preliminary image processing on the initial image. Further, input the preliminarily processed initial image into the artificial intelligence / machine learning model, and automatically generate the text marking content corresponding to each candidate detection box based on the category of the main object included in each candidate detection box.
[0121] S602, according to the text marking content corresponding to each candidate detection box and the category of the main object included in each candidate detection box, determine the priority information of each candidate detection box.
[0122] The text marking content corresponding to the candidate detection box can, to a certain extent, characterize the characteristics of the main object included in the candidate detection box. Combining the text marking content corresponding to the candidate detection box and the category of the main object included in the candidate detection box, determine the priority information of the candidate detection box to ensure as much as possible that the main object included in the candidate detection box with a higher priority is the main object that the user intends to emphasize, that is, the main object to be text-marked.
[0123] Optionally, combining the text marker content corresponding to each candidate detection box, determine the category that the user intends to emphasize from the categories of the main objects included in each candidate detection box, that is, the category of the initial image. The better the performance of the text marker content, the higher the priority of the corresponding candidate detection box. The higher the overlap degree of the categories of the main objects included in different candidate detection boxes, the greater the probability that this category is the category of the initial image.
[0124] Exemplarily, if it is determined that the category of the initial image is food, then the candidate detection box whose main object category is food has a higher priority. Among them, the better the performance of the text marker content, the higher the priority of the candidate detection box, so as to comprehensively determine the priority information of each candidate detection box.
[0125] In this embodiment, a method for determining the priority information of each candidate detection box in the initial image under the one-key marking function is provided. By combining the text marker content corresponding to the candidate detection box and the category of the main object included in the candidate detection box, the determined priority information has higher accuracy, can perceive the user's need to publish graphic and text information as much as possible without the user inputting requirements, is beneficial to subsequent text marking of more appropriate main objects, and optimizes the overall effect of one-key marking.
[0126] In an exemplary embodiment, as Figure 7 shown, the above S601 includes:
[0127] S701, through a multi-modal graphic and text model, extract the image features of each candidate detection box in the initial image and the text features of the template marker content.
[0128] Among them, the multi-modal image model can be a model running on the terminal device or a model running on the server. When the multi-modal graphic and text model runs on the server, the terminal device can transmit the initial image to the server for processing, and then obtain the processing result feedback by the server.
[0129] The multi-modal graphic and text model can include multiple sub-models, which are respectively used to extract the image features of each candidate detection box in the initial image and extract the text features of the template marker content. Among them, the template marker content can be generated by the multi-modal graphic and text model in advance or in real time, or can be pre-input by the application developer based on big data. This embodiment does not limit this.
[0130] The image features of each candidate detection box are actually the image features of the main objects included in each candidate detection box, and are used to characterize visual features including but not limited to the category, position, size, and color of the main object. The text features of the template marker content are actually the text features of different main objects, and are used to characterize various surface or internal semantic features of the main object.
[0131] S702. Perform one-to-one matching on the image features and text features to obtain the graphic-text matching result.
[0132] Perform graphic-text matching on the image features and text features. Specifically, learn the feature representation methods of the image features and text features, calculate the semantic correlation between the image features and text features through similarity, and obtain the graphic-text matching result. The graphic-text matching result includes the text feature with the highest matching degree to the image feature, that is, the template marking content with the highest matching degree to the candidate detection box can be obtained.
[0133] Optionally, for the image features of each candidate detection box, determine a relatively small range of template marking content according to the category of the main object included in the candidate detection box, and then perform one-to-one matching on the image features of the candidate detection box and the text features of the selected template marking content, and determine the text feature with the highest matching degree therefrom to obtain the graphic-text matching result. For example, if the category of the main object included in the candidate detection box is food, then match the image features of the candidate detection box with the text features of the template marking content used to describe food.
[0134] S703. Determine the text marking content corresponding to each candidate detection box from the template marking content according to the graphic-text matching result.
[0135] For each candidate detection box, according to the graphic-text matching result, use the template marking content to which the text feature with the highest matching degree to the image feature of the candidate detection box belongs as the text marking content corresponding to the candidate detection box. Then obtain the text marking content corresponding to all candidate detection boxes.
[0136] In this embodiment, a method for obtaining the text marking content corresponding to the candidate detection box in the one-key marking function is provided. Graphic-text matching is completed through a multi-modal graphic-text model to obtain the text marking content that can accurately describe the main object included in the candidate detection box, reflecting the reliability of intelligent marking and optimizing the overall effect of one-key marking.
[0137] In an exemplary embodiment, as Figure 8 shown, the above S602 includes:
[0138] S801. Determine the category of the initial image according to the text marking content corresponding to each candidate detection box and the category of the main object included in each candidate detection box.
[0139] According to the text marking content corresponding to each candidate detection box, the category of the main object included in each candidate detection box can be accurately refined. For example, if the category of the main object included in the candidate detection box is food, according to the text marking content corresponding to the candidate detection box, determine that the category of the main object included in the candidate detection box is hot pot.
[0140] Further, according to the categories of the main objects included in all candidate detection boxes, determine the category of the initial image, such as taking the category with a higher coincidence degree as the category of the initial image. Exemplarily, there are five candidate detection boxes in the initial image, among which the categories of the main objects included in three candidate detection boxes are food. Specifically, the categories can be determined as hot pot, rice, and fruit according to the text marking content corresponding to the candidate detection boxes. The categories of the main objects included in the other two candidate detection boxes are backpack and water cup. Then, it can be determined that the category of the initial image is food.
[0141] S802. Match the category of the initial image with the categories of the main objects included in each candidate detection box one by one to obtain a category matching result.
[0142] Match the categories of the main objects included in the five candidate detection boxes with the category of the initial image respectively to obtain corresponding category matching results. Among them, the category matching result includes whether there is a match and the similarity of the match, which can be represented by numerical values.
[0143] S803. Determine the priority information of each candidate detection box according to the category matching result.
[0144] Exemplarily, according to the category matching result, it can be determined that three candidate detection boxes whose categories of the main objects included are food have a higher priority than the candidate detection boxes whose categories of the main objects included are backpack or water cup. In addition, different categories correspond to different priorities, and the priority information of all candidate detection boxes is calculated.
[0145] In this embodiment, a method for determining the priority information of candidate detection boxes in the one-key marking function is provided. Combining the text marking content corresponding to the candidate detection boxes and the categories of the main objects included in the candidate detection boxes, it perceives the user's intention, determines the main object that the user is most likely to want to mark from the main objects included in multiple candidate detection boxes, and optimizes the overall effect of one-key marking in a data-based manner with the help of the priority information.
[0146] In an exemplary embodiment, as Figure 9 shown, when the user selects the intelligent marking function, the above S502 includes:
[0147] S901. Determine the confidence information corresponding to each candidate detection box according to the category of the main object included in each candidate detection box.
[0148] It can be understood that when the initial image is detected by the basic detection model, it is actually a prediction of the category of the main object included in each candidate detection box in the initial image. The prediction result includes one or more possible categories, as well as the confidence information corresponding to each category. Furthermore, the category with the highest confidence indicated by the confidence information is used as the category of the main object included in the candidate detection box. Among them, the confidence information is used to characterize the accuracy of the prediction result.
[0149] For each candidate detection box, the confidence information corresponding to the category of the main object included in the candidate detection box is used as the confidence information corresponding to the candidate detection box.
[0150] S902. Determine the priority information of each candidate detection box according to the confidence information and position information corresponding to each candidate detection box.
[0151] The higher the confidence indicated by the confidence information corresponding to the candidate detection box, the higher the priority of the candidate detection box. The closer the position information corresponding to the candidate detection box is to the center of the initial image, the higher the priority of the candidate detection box.
[0152] Optionally, first, according to the position information corresponding to each candidate detection box, perform a preliminary sorting on each candidate detection box, and then, according to the confidence information corresponding to each candidate detection box, correct the result of the preliminary sorting to obtain the priority information of each candidate detection box.
[0153] Optionally, for each candidate detection box, perform a weighted processing on the confidence information and position information corresponding to the candidate detection box according to the confidence weight and position weight to obtain the priority information of the candidate detection box. Assign corresponding weights to the confidence information and position information, and then perform a weighted calculation to obtain the priority information of the candidate detection box. Exemplarily, both the confidence information and the position information are percentages between 0 and 1. The position information can be calculated through the center point coordinates. The closer the center point coordinates of the candidate detection box are to the center point coordinates of the initial image, the larger the value of the position information of the candidate detection box.
[0154] In this embodiment, a method for determining the priority information of candidate detection boxes in the intelligent marking function is provided. By combining the confidence information and position information corresponding to the candidate detection boxes, it perceives the user's intention, determines the main object that the user is most likely to want to mark from the main objects included in multiple candidate detection boxes, so as to provide a template for text marking content with higher user satisfaction, and optimize the overall effect of intelligent marking in a data-based manner by means of the priority information.
[0155] In an exemplary embodiment, as Figure 10 shown, the above text marking method further includes:
[0156] S1001, generate an updated text selection control according to the text marking content corresponding to the target detection box.
[0157] S1002, in response to a trigger operation on the updated text selection control, determine the target text marking content from the text marking content corresponding to the target detection box.
[0158] When the user selects the intelligent marking function, it is also possible to generate a corresponding text selection control according to the text marking content corresponding to the selected candidate detection box, and display the initial image and the generated text selection control on the text marking page, so as to provide it to the user for selection in an intuitive visual way. Furthermore, after the terminal device senses the operation on the text marking content, it determines the candidate detection box corresponding to the text marking content selected by the user as the target detection box, and determines the target text marking content according to the text marking content selected by the user.
[0159] Optionally, the text marking content selected by the user can be directly determined as the target text marking content, or a part of the text marking content can be selected from the text marking content selected by the user as the target text marking content. Exemplarily, the terminal device responds to a trigger operation on the text selection controls corresponding to five different text marking contents, and selects three from the five different text marking contents as the target text marking contents.
[0160] S1003, render the target text marking content to obtain the initial image marked based on the text marking content.
[0161] Render the target text marking content to obtain the marked initial image.
[0162] In this embodiment, it is clarified that under the intelligent marking function, a template of text marking content with higher user satisfaction can be provided to the user first, and then the target text marking content can be determined from it based on the user's selection to complete the rendering. Compared with manually inputting text for marking, the intelligent marking function can improve the user experience and meet the personalized needs of different users.
[0163] In an exemplary embodiment, the above text marking method further includes:
[0164] Display an image editing page; the image editing page displays the initial image, and the image editing page includes a text editing control.
[0165] As Figure 11 shown, before the terminal device displays the text marking page, display the image editing page. The terminal device senses a trigger operation on the new image control in the image editing page, and provides the user with the function of importing the initial image from the album or taking the initial image by using the camera function, so as to display the initial image on the image editing page.
[0166] The text editing control can be represented by an icon of "text".
[0167] Optionally, the image editing page further includes controls for other image editing functions, such as icons represented as "Effect", "Sticker", and "Crop".
[0168] Optionally, the image editing page further includes a music selection control, which is an icon represented as "Select Music", and can provide the user with the function of importing music from local or network.
[0169] Optionally, the image editing page further includes image quantity information. In the case of selecting multiple initial images for editing, the total quantity of the initial images can be displayed through the image quantity information, as well as which one of the initial images is currently being edited, so as to optimize the image editing experience.
[0170] Correspondingly, the above S201 includes:[[]]
[0171] In response to a trigger operation on the text editing control, display a text marking page.
[0172] In response to a trigger operation on the text editing control, the terminal device switches the current page from the image editing page to the text marking page.
[0173] In this embodiment, a visual process for jumping from the image marking page to the text marking page is provided, which improves the steps of the entire text marking method, integrates the text marking function into the image editing function, and is beneficial to the deployment in practical applications.
[0174] In an exemplary embodiment, when the user selects the one-key marking function, the above S202 includes:[[]]
[0175] In response to a trigger operation on the text generation control, obtain the text marking content corresponding to the target detection box in the initial image; wherein, the target detection box is determined from each candidate detection box according to the priority information, and the priority information is determined according to the text marking content corresponding to each candidate detection box and the category of the main object included in each candidate detection box; render the text marking content to obtain the initial image marked based on the text marking content.
[0176] In response to a trigger operation on the text generation control, the terminal device performs category detection on the initial image to obtain the categories of the main objects included in each candidate detection box in the initial image, and then obtains the text marking content corresponding to each candidate detection box. According to the text marking content corresponding to each candidate detection box and the category of the main object included in each candidate detection box, determine the priority information of each candidate detection box, thereby determine the target detection box from each candidate detection box, obtain the text marking content corresponding to the target detection box and perform rendering, and display it in the initial image on the effect preview page.
[0177] Among them, obtaining the text marking content corresponding to each candidate detection box includes: extracting the image features of each candidate detection box and the text features of the template marking content through a multimodal image-text model, performing one-to-one matching on the image features and the text features to obtain an image-text matching result, and determining the text marking content corresponding to each candidate detection box from the template marking content according to the image-text matching result. Determining the priority information of each candidate detection box includes: determining the category of the initial image according to the text marking content corresponding to each candidate detection box and the category of the main object included in each candidate detection box, performing one-to-one matching on the category of the initial image and the category of the main object included in each candidate detection box to obtain a category matching result, and determining the priority information of each candidate detection box according to the category matching result.
[0178] Optionally, when the terminal device executes the above one-key marking step, a one-key marking prompt message is displayed on the text marking page, which can be expressed as "One-key marking in progress...". The one-key marking prompt message includes a cancel control that can be triggered. Before the display effect preview page is displayed, the user can click the cancel control to cancel the one-key marking step.
[0179] In this embodiment, when the user selects the one-key marking function, the terminal device can execute the corresponding steps at the backend, automatically identify the priority information of each candidate detection box in the initial image, determine the main object included in the target detection box that the user is most likely to want to mark according to the priority information, and then complete the text rendering to obtain the target image, and display the image-text effect after one-key marking to the user.
[0180] In an exemplary embodiment, when the user selects the intelligent marking function, S202 above includes:
[0181] In response to the triggering operation on the text selection control, obtaining the template marking content corresponding to the text selection control; wherein, the text selection control is generated according to the target detection box in the initial image, the target detection box is determined from each candidate detection box according to the priority information, and the priority information is determined according to the confidence information and position information corresponding to each candidate detection box; using the template marking content as the text marking content, and rendering the text marking content to obtain the initial image marked based on the text marking content.
[0182] Before the terminal device responds to the trigger operation of the text selection control, after responding to the trigger operation of the text editing control, the initial image can be subjected to category detection when the initial image is acquired, and the category of the main object contained in each candidate detection frame in the initial image can be obtained, and then the text mark content corresponding to each candidate detection frame can be obtained, and the priority information of each candidate detection frame can be determined according to the text mark content corresponding to each candidate detection frame and the category of the main object contained in each candidate detection frame, so as to determine the target detection frame from each candidate detection frame. The terminal device obtains the text mark content corresponding to the target detection frame and generates the corresponding text selection control.
[0183] Furthermore, in response to the triggering operation on the text selection control, the terminal device obtains the text mark content corresponding to the target detection box and renders it, and displays it in the initial image of the effect preview page.
[0184] Determining the priority information of each candidate detection frame includes: determining the confidence information corresponding to each candidate detection frame according to the category of the main object contained in each candidate detection frame, and determining the priority information of each candidate detection frame according to the confidence information and position information corresponding to each candidate detection frame. Exemplarily, for each candidate detection frame, the confidence information and position information corresponding to the candidate detection frame are weighted according to the confidence weight and the position weight to obtain the priority information of the candidate detection frame.
[0185] Optionally, when the terminal device executes the above-mentioned smart tagging step, only the text tag content selected by the user is rendered, that is, only the text tag content corresponding to the target detection box selected by the user is rendered. The terminal can respond to one trigger operation on the text selection control at a time, or can respond to multiple trigger operations on the text selection control at a time, and this embodiment does not limit this.
[0186] In this embodiment, when the user selects the smart tagging function, the terminal device can execute corresponding steps in the back end, automatically identify the priority information of each candidate detection box in the initial image, and determine the main object contained in the target detection box that the user is most likely to want to mark based on the priority information, and then generate a corresponding text selection control and display it on the text tagging page until the user's triggering operation on the text selection control is sensed, the text rendering is completed to obtain the target image, and the graphic effect after smart tagging is displayed to the user.
[0187] In an exemplary embodiment, the text marking method further includes:
[0188] In response to a triggering operation of an update control of the text markup content, an updated effect preview page is displayed; the updated effect preview page displays an updated initial image, and the update control includes at least one of an edit update control, a mirror update control, and a delete update control.
[0189] As Figure 12 shown, in the effect preview page, the user can click on the text marker content on the upper layer of the initial image to display the update controls for the text marker content, namely the edit update control, the mirror update control, and the delete update control. Among them, the edit update control is used to provide an interface for triggering the re-editing operation of the text marker content, the mirror update control is used to trigger an interface for mirroring the text marker content, and the delete update control is used to trigger an interface for deleting the generated text marker content.
[0190] Optionally, the terminal device can respond to the trigger operation on the update control of the text marker content and jump to the image editing page to display the updated target image in the image editing page.
[0191] In this embodiment, the user can also modify the generated text marker content in a visual manner, improving the text marking process and optimizing the user experience.
[0192] In an exemplary embodiment, a text marking method is provided, including:
[0193] The terminal device responds to the trigger operation of the user on the text input control in the text marking page and executes the following custom marking process: obtaining the text marker content input through the text input box corresponding to the text input control, and rendering the text marker content to obtain the target image corresponding to the initial image. Further, the terminal device displays the target image after text marking in the effect preview page, and the user can make corresponding adjustments to the custom-generated text marker content.
[0194] In an exemplary embodiment, a text marking method is provided, including:
[0195] In response to a user's triggering operation on a text generation control in a text marking page, the terminal device executes the following one-key marking process: performing category detection on an initial image to be marked to obtain the categories of the main objects included in each candidate detection box in the initial image; wherein each candidate detection box contains one main object; extracting the image features of each candidate detection box in the initial image and the text features of the template marking content through a multimodal image-text model; performing one-to-one matching on the image features and the text features to obtain an image-text matching result; determining the text marking content corresponding to each candidate detection box from the template marking content according to the image-text matching result. Determining the category of the initial image according to the text marking content corresponding to each candidate detection box and the category of the main object included in each candidate detection box; performing one-to-one matching on the category of the initial image and the categories of the main objects included in each candidate detection box to obtain a category matching result; determining the priority information of each candidate detection box according to the category matching result. Determining a target detection box from each candidate detection box according to the priority information; rendering the text marking content corresponding to the target detection box to obtain a target image corresponding to the initial image. Further, the terminal device displays the target image after text marking in an effect preview page, and the user can make corresponding adjustments to the text marking content generated by one key.
[0196] In an exemplary embodiment, a text marking method is provided, including:
[0197] In response to a user's triggering operation on a text editing control in an image editing page, the terminal device obtains an initial image to be marked and executes the following intelligent marking process: performing category detection on the initial image to be marked to obtain the categories of the main objects included in each candidate detection box in the initial image; wherein each candidate detection box contains one main object; determining the confidence information corresponding to each candidate detection box according to the category of the main object included in each candidate detection box; performing weighted processing on the confidence information and the position information corresponding to each candidate detection box according to the confidence weight and the position weight to obtain the priority information of each candidate detection box. Determining a target detection box from each candidate detection box according to the priority information; generating a text selection control for the text marking content corresponding to the target detection box. Further, in response to a user's triggering operation on the text selection control in the text marking page, the terminal device determines the target text marking content from the text marking content corresponding to the target detection box; rendering the target text marking content to obtain a target image corresponding to the initial image. Displaying the target image after text marking in an effect preview page, and the user can make corresponding adjustments to the text marking content generated intelligently.
[0198] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0199] Based on the same inventive concept, an embodiment of the present application further provides a text marking device for implementing the above-mentioned text marking method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the following text marking devices can refer to the limitations on the text marking method in the above, and will not be repeated here.
[0200] In an exemplary embodiment, as Figure 13 shown, a text marking device is provided, including:
[0201] A marked page display module 10 for displaying a text marking page; the text marking page displays an initial image to be marked, and the text marking page includes text marking controls; wherein, the text marking controls include at least one of a text input control, a text generation control, and a text selection control.
[0202] An acquisition module 20 for acquiring text marking content in response to a trigger operation on the text marking control.
[0203] An image generation module 30 for generating a target image corresponding to the initial image, the target image including the text marking content.
[0204] In an exemplary embodiment, the above acquisition module 20 includes:
[0205] A content acquisition sub-module for acquiring the text marking content input through the text input control in response to a trigger operation on the text input control.
[0206] A text marking sub-module for rendering the text marking content to obtain the initial image marked based on the text marking content.
[0207] In an exemplary embodiment, on the basis of Figure 13 , as Figure 14 shown, the above acquisition module 20 includes:
[0208] The category detection sub-module 21 is used to perform category detection on the initial image to be marked, and obtain the category of the main object included in each candidate detection box in the initial image; wherein, each candidate detection box includes one main object.
[0209] The priority determination sub-module 22 is used to determine the priority information of each candidate detection box according to the category of the main object included in each candidate detection box.
[0210] The detection box determination sub-module 23 is used to determine the target detection box from each candidate detection box according to the priority information.
[0211] The marking rendering sub-module 24 is used to render the text marking content corresponding to the target detection box, and obtain the initial image marked based on the text marking content.
[0212] In an exemplary embodiment, the above-mentioned priority determination sub-module 22 includes:
[0213] The marking content acquisition unit is used to acquire the text marking content corresponding to each candidate detection box in the initial image.
[0214] The first priority determination unit is used to determine the priority information of each candidate detection box according to the text marking content corresponding to each candidate detection box and the category of the main object included in each candidate detection box.
[0215] In an exemplary embodiment, the above-mentioned marking content acquisition unit includes:
[0216] The feature extraction sub-unit is used to extract the image features of each candidate detection box in the initial image and the text features of the template marking content through a multi-modal graphic and text model.
[0217] The graphic and text matching sub-unit is used to perform one-to-one matching on the image features and text features to obtain a graphic and text matching result.
[0218] The marking content acquisition sub-unit is used to determine the text marking content corresponding to each candidate detection box from the template marking content according to the graphic and text matching result.
[0219] In an exemplary embodiment, the above-mentioned first priority determination unit includes:
[0220] The category determination sub-unit is used to determine the category of the initial image according to the text marking content corresponding to each candidate detection box and the category of the main object included in each candidate detection box.
[0221] The category matching sub-unit is used to perform one-to-one matching on the category of the initial image and the category of the main object included in each candidate detection box to obtain a category matching result.
[0222] A first priority determination subunit, configured to determine the priority information of each candidate detection box according to the category matching result.
[0223] In an exemplary embodiment, the above-mentioned priority determination sub-module 22 includes:
[0224] A confidence determination unit, configured to determine the confidence information corresponding to each candidate detection box according to the category of the main object included in each candidate detection box.
[0225] A second priority determination unit, configured to determine the priority information of each candidate detection box according to the confidence information and position information corresponding to each candidate detection box.
[0226] In an exemplary embodiment, the above-mentioned second priority determination unit includes:
[0227] A second priority determination subunit, configured to, for each candidate detection box, perform weighted processing on the confidence information and position information corresponding to the candidate detection box according to the confidence weight and position weight, to obtain the priority information of the candidate detection box.
[0228] In an exemplary embodiment, the above-mentioned marker rendering sub-module 24 includes:
[0229] A control update unit, configured to generate an updated text selection control according to the text marker content corresponding to the target detection box.
[0230] A content screening unit, configured to, in response to a trigger operation on the updated text selection control, determine target text marker content from the text marker content corresponding to the target detection box.
[0231] A marker rendering unit, configured to render the target text marker content to obtain an initial image marked based on the text marker content.
[0232] In an exemplary embodiment, the above-mentioned text marking device further includes:
[0233] An image page display module, configured to display an image editing page; the image editing page displays an initial image, and the image editing page includes a text editing control.
[0234] Correspondingly, the above-mentioned marker page display module 10 is specifically configured to display a text marking page in response to a trigger operation on the text editing control.
[0235] In an exemplary embodiment, the above-mentioned text marking device further includes:
[0236] A page update module is configured to display a preview page of the updated effect in response to a triggering operation on an update control for text marked content; the preview page of the updated effect displays an updated initial image, and the update control includes at least one of an edit update control, a mirror update control, and a delete update control.
[0237] Each module in the above text marking device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.
[0238] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 15 shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communicating with external terminals in a wired or wireless manner. The wireless manner can be achieved through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. The computer program, when executed by the processor, implements a text marking method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0239] Those skilled in the art can understand that Figure 15 the structure shown in
[0240] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the above text marking method are implemented.
[0241] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above text marking method are implemented.
[0242] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps of the above text marking method are implemented.
[0243] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., and are not limited thereto.
[0244] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in the present application.
[0245] The above embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A text marking method, characterized in that, The method includes: Displaying a text marking page; the text marking page displays an initial image to be marked, and the text marking page includes text marking controls; wherein, the text marking controls include at least one of a text input control, a text generation control, and a text selection control; In response to a triggering operation on the text marking control, obtaining text marking content; Generating a target image corresponding to the initial image, where the target image includes the text marking content.
2. The method according to claim 1, wherein The obtaining text marking content in response to a triggering operation on the text marking control includes: In response to a triggering operation on the text input control, obtaining text marking content input through the text input control; Rendering the text marking content to obtain the initial image marked based on the text marking content.
3. The method according to claim 1, characterized in that, The obtaining text marking content in response to a triggering operation on the text marking control includes: In response to a triggering operation on the text marking control, performing category detection on the initial image to be marked to obtain the categories of the main objects included in each candidate detection box in the initial image; wherein each candidate detection box includes one main object; Determining priority information for each candidate detection box according to the categories of the main objects included in each candidate detection box; Determining a target detection box from each candidate detection box according to the priority information; Rendering the text marking content corresponding to the target detection box to obtain the initial image marked based on the text marking content.
4. The method according to claim 3, characterized in that When the text marking control is a text generation control, the determining priority information for each candidate detection box according to the categories of the main objects included in each candidate detection box includes: Obtaining the text marking content corresponding to each candidate detection box in the initial image; Determining priority information for each candidate detection box according to the text marking content corresponding to each candidate detection box and the categories of the main objects included in each candidate detection box.
5. The method according to claim 4, wherein The obtaining the text marking content corresponding to each candidate detection box in the initial image includes: Extracting the image features of each candidate detection box and the text features of the template marking content in the initial image through a multimodal image-text model; Performing one-to-one matching on the image features and the text features to obtain an image-text matching result; Determining the text marking content corresponding to each candidate detection box from the template marking content according to the image-text matching result.
6. The method according to claim 4, wherein The determining priority information for each candidate detection box according to the text marking content corresponding to each candidate detection box and the categories of the main objects included in each candidate detection box includes: Determining the category of the initial image according to the text marking content corresponding to each candidate detection box and the categories of the main objects included in each candidate detection box; Performing one-to-one matching on the category of the initial image and the categories of the main objects included in each candidate detection box to obtain a category matching result; Determining priority information for each candidate detection box according to the category matching result.
7. The method according to claim 3, wherein When the text marking control is a text selection control, the determining priority information for each candidate detection box according to the categories of the main objects included in each candidate detection box includes: Determine the confidence information corresponding to each candidate detection box according to the category of the main object included in the candidate detection box; Determine the priority information of each candidate detection box according to the confidence information and position information corresponding to each candidate detection box.
8. The method according to claim 7, wherein The determining the priority information of each candidate detection box according to the confidence information and position information corresponding to each candidate detection box includes: For each candidate detection box, perform weighted processing on the confidence information and position information corresponding to the candidate detection box according to the confidence weight and position weight to obtain the priority information of the candidate detection box.
9. The method according to claim 8, wherein The rendering the text marking content corresponding to the target detection box to obtain the initial image marked based on the text marking content includes: Generate an updated text selection control according to the text marking content corresponding to the target detection box; In response to the triggering operation on the updated text selection control, determine the target text marking content from the text marking content corresponding to the target detection box; Render the target text marking content to obtain the initial image marked based on the text marking content.
10. The method according to claim 1, characterized in that, Before displaying the text marking page, the method further includes: Display an image editing page; the image editing page displays the initial image, and the image editing page includes a text editing control; Correspondingly, the displaying the text marking page includes: In response to the triggering operation on the text editing control, display the text marking page.
11. The method according to claim 1, wherein Before generating the target image corresponding to the initial image, the method further includes: In response to the triggering operation on the update control of the text marking content, display an updated effect preview page; the updated effect preview page displays the updated initial image, and the update control includes at least one of an editing update control, a mirror update control, and a deletion update control.
12. A text marking device, characterized in that, The apparatus includes: A marking page display module, configured to display a text marking page; the text marking page displays an initial image to be marked, and the text marking page includes a text marking control; wherein, the text marking control includes at least one of a text input control, a text generation control, and a text selection control; An obtaining module, configured to obtain text marking content in response to the triggering operation on the text marking control; An image generation module, configured to generate a target image corresponding to the initial image, and the target image includes the text marking content.
13. The device according to claim 12, characterized in that, The obtaining module includes: A category detection unit, configured to perform category detection on the initial image to be marked in response to the triggering operation on the text marking control, and obtain the category of the main object included in each candidate detection box in the initial image; wherein, each candidate detection box includes one main object; A priority determination unit, configured to determine the priority information of each candidate detection box according to the category of the main object included in each candidate detection box; A detection box determination unit, configured to determine a target detection box from each candidate detection box according to the priority information; A marking rendering unit, configured to render the text marking content corresponding to the target detection box to obtain the initial image marked based on the text marking content.
14. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-11.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-11.
16. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-11.