A method, related device, equipment, and storage medium for image processing

By identifying contents of the pictures of the terminal device and generating structure files, the problem of insufficient storage space of the terminal device is solved, and the storage space optimization and information retention are achieved.

CN115116078BActive Publication Date: 2025-07-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210642570.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-08
Publication Date
2025-07-22
Estimated Expiration
2042-06-08

AI Technical Summary

Technical Problem

The problem of insufficient storage space for terminal devices, especially because a large number of text-type images take up too much storage space, deleting pictures in existing solutions is time-consuming and laborious and may delete important information.

Method used

By recognizing the original image, a structure file is generated. The file contains text information and stores the structure file to replace the original image to reduce the storage space.

Benefits of technology

It effectively alleviates the problem of excessive storage space in pictures, reduces storage space, and avoids the loss of important information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115116078B_ABST
    Figure CN115116078B_ABST
Patent Text Reader

Abstract

This application discloses a method for image processing, and the application scenarios at least include various terminals, such as mobile phones, computers, vehicle-mounted terminals, etc. The method provided by this application includes: obtaining an original image to be processed; performing content recognition on the original image to obtain the image type of the original image; if the image type of the original image meets the format conversion condition, generating a structure file corresponding to the original image, the structure file includes at least one text information, and at least one text information is the text content recognized based on the original image; storing the structure file, and the structure file is used to generate a target image. This application also provides related devices, equipment, and storage media. This application generates and stores a structure file based on the original image. Compared with the original image, the structure file occupies less storage space, thereby effectively alleviating the problem of excessive storage space occupied by images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular, to a method for image processing, related devices, equipment, and storage media. Background Art

[0002] Nowadays, social communication applications have become essential daily software for users. Users can conduct a large amount of social interactions through terminal devices such as mobile phones or computers, and a large amount of messages are generated. A large part of these messages are pictures, and with the accumulation of a large number of pictures, it brings a relatively large storage pressure to the storage space of the terminal device.

[0003] Currently, when the storage space of the terminal device is insufficient, users can select a part of the pictures for deletion. These pictures may contain a large number of pictures with text as the main content. Therefore, when users choose to delete, they usually need to preview the entire text content before selectively deleting.

[0004] The inventors found that there are at least the following problems in the existing solutions. However, although directly deleting pictures can free up a certain amount of storage space, the deletion process is time-consuming and laborious, and may also delete important information. Therefore, how to optimize the storage of a large number of pictures and avoid occupying a large amount of storage space is an urgent problem to be solved. Summary of the Invention

[0005] Embodiments of this application provide a method for image processing, related devices, equipment, and storage media. This application generates and stores a structure file based on the original picture. Compared with the original picture, the structure file occupies less storage space, thereby effectively alleviating the problem of excessive storage space occupied by pictures.

[0006] In view of this, on the one hand, this application provides a method for image processing, including:

[0007] Obtain the original picture to be processed;

[0008] Perform content recognition on the original picture to obtain the picture type of the original picture;

[0009] If the picture type of the original picture meets the format conversion condition, generate a structure file corresponding to the original picture, where the structure file includes at least one text information, and at least one text information is the text content recognized from the original picture;

[0010] Store the structure file, where the structure file is used to generate the target picture.

[0011] On the other hand, this application provides an image processing device, including:

[0012] An acquisition module, configured to acquire an original picture to be processed;

[0013] The acquisition module is further configured to perform content recognition on the original picture to obtain the picture type of the original picture;

[0014] A generation module, configured to generate a structure file corresponding to the original picture if the picture type of the original picture meets the format conversion condition, where the structure file includes at least one piece of text information, and the at least one piece of text information is the text content recognized based on the original picture;

[0015] A storage module, configured to store the structure file, where the structure file is used to generate a target picture.

[0016] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0017] The acquisition module is specifically configured to call a target detection model to perform content recognition on the original picture to obtain a picture recognition result;

[0018] If the picture recognition result indicates that the original picture includes T text regions, determine the proportion of the T text regions in the original picture, where T is an integer greater than or equal to 1;

[0019] If the proportion of the T text regions in the original picture is greater than or equal to a first proportion threshold, determine that the picture type of the original picture is a text picture type, where the text picture type meets the format conversion condition.

[0020] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0021] The generation module is specifically configured to perform optical character recognition (OCR) processing on each of the T text regions to obtain the text information corresponding to the text region;

[0022] Generate a structure file corresponding to the original picture according to the text information corresponding to each of the T text regions.

[0023] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the structure file includes a preset field set, and the preset field set includes a content field;

[0024] The generation module is specifically configured to determine the text information corresponding to the content field according to the text information corresponding to each of the T text regions;

[0025] If the preset field set further includes a region position field, determine the region position information corresponding to the region position field according to the T text regions;

[0026] If the preset field set further includes a font size field, determine the font size information corresponding to the font size field according to the text information corresponding to each of the T text regions.

[0027] If the preset field set further includes a region width field, determine the region width information corresponding to the region width field according to the T text regions.

[0028] If the preset field set further includes a font color field, determine the font color information corresponding to the font color field according to the text information corresponding to each of the T text regions.

[0029] If the preset field set further includes a font bold field, determine the font bold information corresponding to the font bold field according to the text information corresponding to each of the T text regions.

[0030] If the preset field set further includes a font italic field, determine the font italic information corresponding to the font italic field according to the text information corresponding to each of the T text regions.

[0031] If the preset field set further includes a font type field, determine the font type information corresponding to the font type field according to the text information corresponding to each of the T text regions.

[0032] If the preset field set further includes a picture type field, determine the text picture type corresponding to the picture type field.

[0033] Generate a structure file corresponding to the original picture according to the text information and at least one of the region position information, font size information, region width information, font color information, font bold information, font italic information, font type information, and text picture type.

[0034] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the picture processing device further includes a processing module.

[0035] The obtaining module is further configured to obtain a target structure file.

[0036] The processing module is configured to, if the target structure file includes a text picture type, perform rendering processing according to the target structure file to obtain a first target picture, where at least one text segment is displayed on the first target picture.

[0037] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0038] An acquisition module, specifically used to perform content matching on the original image using a first template image to obtain a content matching result;

[0039] If the content matching result indicates a successful match, determine that the image type of the original image is a session image type, where the session image type meets the format conversion conditions.

[0040] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0041] A generation module, specifically used to perform optical character recognition (OCR) processing on the original image to obtain M text messages, where M is an integer greater than or equal to 1;

[0042] Perform content matching on the original image using a second template image to obtain N session areas, where each session area includes a text message, and N is an integer greater than or equal to 1;

[0043] According to the N session areas, obtain the avatar corresponding to each session area;

[0044] Generate a structure file corresponding to the original image according to the M text messages and the avatar corresponding to each session area.

[0045] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0046] The generation module is specifically used to, for each of the N session areas, obtain the avatar corresponding to the session area according to a preset offset;

[0047] The processing module is further used to perform a hash calculation on the avatar corresponding to each session area to obtain the hash information to be matched for each avatar;

[0048] The processing module is further used to, if the hash information to be matched for the avatar matches the avatar hash information stored locally in the terminal device, use the hash information to be matched for the avatar as the avatar hash information;

[0049] The storage module is further used to, if the hash information to be matched for the avatar does not match the avatar hash information stored locally in the terminal device, store the avatar and use the hash information to be matched for the avatar as the avatar hash information.

[0050] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the structure file includes a preset field set, and the preset field set includes a message field and an avatar field;

[0051] The generation module is specifically used to determine the N text messages corresponding to the message field according to the M text messages;

[0052] Determine N avatar hash information corresponding to the avatar field according to the avatars corresponding to each session area;

[0053] If the preset field set further includes a sender field, determine N sender information corresponding to the sender field according to the positional relationship between each session area and the corresponding avatar;

[0054] If the preset field set further includes a picture type field, determine the session picture type according to the picture type field;

[0055] If the preset field set further includes an object name field, determine the object name information corresponding to the object name field according to M text messages;

[0056] If the preset field set further includes an announcement content field, determine the announcement information corresponding to the announcement content field according to M text messages;

[0057] If the preset field set further includes a voice message field, determine the voice type information corresponding to the voice message field according to N text messages;

[0058] Generate a structure file corresponding to the original picture according to the N text messages and the N avatar hash information, and according to at least one of the N sender information, the session picture type, the object name information, the announcement information and the voice type information.

[0059] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0060] The acquisition module is further configured to acquire a target structure file;

[0061] The processing module is further configured to, if the target structure file includes a session picture type, perform rendering processing according to the target structure file to obtain a second target picture, where the second target picture displays avatars, text information in the session area, and object name information.

[0062] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0063] The acquisition module is specifically configured to call a target detection model to perform content recognition on the original picture to obtain a picture recognition result;

[0064] If the picture recognition result indicates that the original picture includes an information code area, determine the proportion of the information code area in the original picture;

[0065] If the proportion of the information code area in the original picture is greater than or equal to a second proportion threshold, determine that the picture type of the original picture is an information code picture type, where the information code picture type meets the format conversion condition.

[0066] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0067] The generation module is specifically configured to call an information code decoder to decode the information code in the information code area to obtain text information, where the text information is the link address corresponding to the information code;

[0068] Generate a structure file corresponding to the original picture according to the text information.

[0069] In a possible design, in another implementation of another aspect of the embodiments of the present application, the structure file includes a preset field set, and the preset field set includes a content field;

[0070] The generation module is specifically configured to determine the text information corresponding to the content field according to the text information;

[0071] If the preset field set further includes a picture type field, determine the information code picture type corresponding to the picture type field;

[0072] If the preset field set further includes a region width field, determine the region width information corresponding to the region width field according to the information code area;

[0073] If the preset field set further includes a region height field, determine the region height information corresponding to the region height field according to the information code area;

[0074] Generate a structure file corresponding to the original picture according to the text information and at least one of the information code picture type, region width information, and region height information.

[0075] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0076] The acquisition module is further configured to acquire a target structure file;

[0077] The processing module is further configured to, if the target structure file includes an information code picture type, perform rendering processing according to the target structure file to obtain a third target picture, where the third target picture displays an information code.

[0078] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0079] The acquisition module is specifically configured to display a space optimization message and a space optimization control, where the space optimization message is used to prompt the available space to be released;

[0080] In response to a selection operation on the space optimization control, acquire the original picture to be processed;

[0081] The processing module is further configured to delete the original picture locally from the terminal device and upload the original picture to the cloud server or other devices.

[0082] In a possible design, in another implementation manner of another aspect of this application embodiment,

[0083] The processing module is further configured to, after storing the structure file, in response to a sending operation for the target picture, send the target picture to the target terminal device;

[0084] Or,

[0085] The processing module is further configured to, after storing the structure file, in response to a sending operation for the target picture, send the structure file to the target terminal device, so that the target picture is obtained through rendering processing according to the target structure file.

[0086] Another aspect of this application provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the methods in the above aspects are implemented.

[0087] Another aspect of this application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the methods in the above aspects are implemented.

[0088] Another aspect of this application provides a computer program product, including a computer program. When the computer program is executed by a processor, the methods in the above aspects are implemented.

[0089] It can be seen from the above technical solutions that the embodiments of this application have the following advantages:

[0090] In the embodiments of this application, a method for picture processing is provided. First, the original picture to be processed is obtained. Then, content recognition is performed on the original picture to obtain the picture type of the original picture. If the picture type of the original picture meets the format conversion condition, a structure file corresponding to the original picture is generated. The structure file here includes text information, and the text information is the text content recognized from the original picture. Based on this, the terminal device stores the structure file. Through the above method, format conversion is performed on the original picture containing text information, that is, the original picture is recognized and relevant text information is extracted, and then the structure file is generated based on the text information. Thus, the terminal device only needs to store the structure file locally. It can be seen that the structure file mainly records text information. Therefore, compared with the original picture, the structure file occupies less storage space, thereby effectively alleviating the problem of excessive storage space occupied by pictures. Description of the Drawings

[0091] Figure 1 A schematic diagram of text-type pictures and non-text-type pictures in an embodiment of the present application;

[0092] Figure 2 A schematic architecture diagram of a picture processing system in an embodiment of the present application;

[0093] Figure 3 A schematic flowchart of a picture processing method in an embodiment of the present application;

[0094] Figure 4 A schematic diagram of identifying a text area based on an original image in an embodiment of the present application;

[0095] Figure 5 A schematic diagram of an original picture and a first target picture in an embodiment of the present application;

[0096] Figure 6 Another schematic diagram of an original picture and a first target picture in an embodiment of the present application;

[0097] Figure 7 A schematic diagram of a first template picture in an embodiment of the present application;

[0098] Figure 8 A schematic diagram of implementing template matching for an original picture using a first template picture in an embodiment of the present application;

[0099] Figure 9 A schematic diagram of a second template picture in an embodiment of the present application;

[0100] Figure 10 A schematic diagram of identifying a session area based on an original image in an embodiment of the present application;

[0101] Figure 11 A schematic diagram of an original picture and a second target picture in an embodiment of the present application;

[0102] Figure 12 Another schematic diagram of an original picture and a second target picture in an embodiment of the present application;

[0103] Figure 13 A schematic diagram of identifying an information code area based on an original image in an embodiment of the present application;

[0104] Figure 14 A schematic diagram of an original picture and a third target picture in an embodiment of the present application;

[0105] Figure 15 A schematic interface diagram of storing session pictures in an IM application in an embodiment of the present application;

[0106] Figure 16It is a schematic diagram of an interface for storing pictures in the album application in the embodiment of the present application;

[0107] Figure 17 It is a schematic diagram for migrating the original picture in the embodiment of the present application;

[0108] Figure 18 It is another schematic diagram for migrating the original picture in the embodiment of the present application;

[0109] Figure 19 It is another schematic diagram for migrating the original picture in the embodiment of the present application;

[0110] Figure 20 It is a schematic diagram of an interface for transmitting the target picture in the embodiment of the present application;

[0111] Figure 21 It is a schematic diagram of a picture processing device in the embodiment of the present application;

[0112] Figure 22 It is a schematic diagram of the structure of a terminal device in the embodiment of the present application. Detailed implementation manners

[0113] The embodiment of the present application provides a method, related device, equipment and storage medium for picture processing. The present application generates and stores a structure file based on the original picture. Compared with the original picture, the structure file occupies less storage space, thereby effectively alleviating the problem of excessive storage space occupied by pictures.

[0114] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or equipment including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.

[0115] With the advent of the high-speed communication era, while bringing us convenient experiences, the accompanying issue is the sharing and storage of a vast amount of data. In many cases, a lot of text-based pictures are stored in the user's terminal devices (such as mobile phones, tablets, etc.). These pictures often occupy a large amount of storage space and can easily cause insufficient storage space on the terminal device. As the name implies, text-based pictures directly display text content or can be processed into text content. For ease of understanding, please refer to Figure 1 , Figure 1 which is a schematic diagram of text-based pictures and non-text-based pictures in an embodiment of this application. Figure 1 In (A) of [the figure], it shows a text-based picture, where the text-based picture includes text content. Figure 1 In (B) of [the figure], it shows a non-text-based picture, where the non-text-based picture does not include text content.

[0116] To alleviate the problem of text-based pictures occupying too much storage space. This application proposes a method for picture processing, which is applied to Figure 2 the picture processing system shown as follows. As shown in the figure, the picture processing system includes a terminal device 110 and a server 120, and the client is deployed on the terminal device. Among them, the client can run on the terminal device in the form of a browser or in the form of an independent application (APP), etc. Regarding the specific display form of the client, it is not limited here. The server involved in this application can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal device can be a mobile phone, a tablet, a laptop, a handheld computer, a personal computer, a smart TV, a smart watch, a vehicle-mounted device, a wearable device, etc., but is not limited thereto. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not limit this here. The number of servers and terminal devices is also not limited. The solution provided by this application can be completed independently by the terminal device, independently by the server, or jointly by the terminal device and the server. Regarding this, this application does not make specific limitations.

[0117] Exemplarily, after the terminal device recognizes a text-based picture, it generates a structure file corresponding to the text-based picture. Based on this, on the one hand, the structure file can be stored locally on the terminal device, and on the other hand, the text-based picture can be selected to be uploaded to the cloud server for storage.

[0118] It is understandable that the image processing method provided in this application involves technologies such as computer vision (CV) in artificial intelligence (AI). Among them, CV technology is a science that studies how to enable machines to "see". Further, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition and measurement on targets, and further perform graphic processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, CV studies related theories and technologies and attempts to establish an AI system that can obtain information from images or multi-dimensional data. CV technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0119] Among them, AI uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, AI is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0120] AI technology is a comprehensive discipline with a wide range of fields involved, including both hardware-level technologies and software-level technologies. AI basic technologies generally include technologies such as sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. AI software technologies mainly include several major directions such as CV technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0121] In view of the fact that this application involves some terms related to professional fields, for the sake of easy understanding, explanations will be given below.

[0122] (1) Instant messaging (IM) application: It is a terminal service that allows two or more people to instantaneously transmit text messages, files, voice, and video communication, etc. over the network.

[0123] (2) OCR: It refers to the process of analyzing and recognizing images of text materials to obtain text and layout information.

[0124] (3) Byte: 8 binary digits form 1 byte, which is the basic measurement unit of computer storage space. 1 kilobyte (KB) represents 1024 bytes.

[0125] (4) Template matching: It refers to the technology of finding the most matching or similar part in a picture with another template picture.

[0126] (5) OpenCV: It is a cross-platform computer vision library distributed under the Berkeley Software Distribution (BSD) license and can run on Linux operating system, Windows operating system, Android operating system, and Mac OS operating system.

[0127] (6) Hash information: Any length of input (also known as pre-mapping) is transformed into a fixed-length output through a hashing algorithm, and this output is the hash information.

[0128] Combined with the above introduction, the solution provided in the embodiments of this application involves technologies such as CV in AI, and is specifically described through the following embodiments. Please refer to Figure 3 , the image processing method in the embodiments of this application can be executed by a terminal device or a server. The image processing method provided in this application includes:

[0129] 210. Obtain the original image to be processed;

[0130] In one or more embodiments, the original image to be processed can be obtained from the images stored locally in the terminal device, or from the images cached in the terminal device, or from the images stored in other devices, which is not limited here. Among them, the original image can be a picture obtained by the user's screenshot, or a picture downloaded by the user, or a picture taken by the user, etc.

[0131] Specifically, taking the example of obtaining the original image by taking a screenshot in an IM application, for example, the user searches for news, keywords, etc. through the IM application. Based on this, a search page is displayed on the IM application interface. After the user triggers the screenshot, an original image is obtained. For example, the user chats with other users through the IM. Based on this, a conversation page is displayed on the IM application interface. After the user triggers the screenshot, an original image is obtained. For example, the user opens the personal identification code through the IM. Based on this, an information code page is displayed on the IM application interface. After the user triggers the screenshot, an original image is obtained.

[0132] 220. Perform content recognition on the original image to obtain the image type of the original image;

[0133] In one or more embodiments, perform content recognition on the original image to obtain the image recognition result of the original image, and the image type of the original image can be determined according to the image recognition result. The following will introduce several ways of content recognition:

[0134] Method 1: Directly call the object detection model to perform content recognition on the original image.

[0135] Method 2: Directly use the image template to perform content recognition on the original image.

[0136] Method 3: Call the object detection model and use the image template to perform content recognition on the original image.

[0137] 230. If the image type of the original image meets the format conversion condition, generate the structure file corresponding to the original image, where the structure file includes at least one text information, and at least one text information is the text content recognized based on the original image;

[0138] In one or more embodiments, if the image type of the original image belongs to the image type of the convertible format, it means that the original image meets the format conversion condition. Thus, further recognition is performed on the original image to obtain at least one text information. Then, a structure is generated based on at least one text information and written into the corresponding structure file. Among them, the structure is a data type in the C language, usually used to represent several related data of different types.

[0139] Specifically, assume that after performing OCR recognition on the original image, a text information is obtained, which is "Rose is a general name for various plants and cultivated flowers in the Rosales order, Rosaceae family, and Rosa genus." That is, the structure can be expressed as:

[0140] {

[0141] "content":"Rose is a general name for various plants and cultivated flowers in the Rosales order, Rosaceae family, and Rosa genus."

[0142] }

[0143] Among them, "content" represents the content field. "Rose is a general name for various plants and cultivated flowers in the Rosales order, Rosaceae family, and Rosa genus." represents a text information.

[0144] 240. Store the structure file, where the structure file is used to generate the target image.

[0145] In one or more embodiments, the generated structure file is stored locally on the terminal device. Based on this, the user can choose to delete the original picture, or choose to save the original picture locally, or directly replace the original picture with the structure file for storage. When the storage space of the terminal device is insufficient, the original picture stored locally can be preferentially deleted. Further, if the original picture needs to be viewed, the structure file is directly used for rendering to generate a target picture. The target picture retains important elements in the original picture (such as text content or information codes, etc.).

[0146] In the embodiments of the present application, a method for picture processing is provided. Through the above method, the original picture containing text information is subjected to format conversion, that is, the original picture is recognized and relevant text information is extracted, and then a structure file is generated based on the text information. Thus, only the structure file needs to be stored locally on the terminal device. It can be seen that the structure file mainly records text information. Therefore, compared with the original picture, the structure file occupies less storage space, thereby effectively alleviating the problem of excessive storage space occupied by pictures.

[0147] Optionally, on the basis of the above Figure 3 corresponding various embodiments, in another optional embodiment provided by the embodiments of the present application, the content of the original picture is recognized to obtain the picture type of the original picture, which may specifically include:

[0148] Call a target detection model to perform content recognition on the original picture to obtain a picture recognition result;

[0149] If the picture recognition result indicates that the original picture includes T text regions, determine the proportion of the T text regions in the original picture, where T is an integer greater than or equal to 1;

[0150] If the proportion of the T text regions in the original picture is greater than or equal to the first proportion threshold, determine that the picture type of the original picture is a text picture type, where the text picture type meets the format conversion conditions.

[0151] In one or more embodiments, a method for identifying a text picture type is introduced. As can be seen from the foregoing embodiments, a target detection model can be directly called to perform content recognition on the original picture. Among them, the input of the target detection model is a picture, and the output is a bounding box and the category label corresponding to the bounding box. In this embodiment, the bounding box is the bounding box corresponding to the text region, and the category label corresponding to the bounding box is "text class".

[0152] Specifically, for ease of understanding, please refer to Figure 4 , Figure 4 which is a schematic diagram for recognizing text regions based on the original image in the embodiments of the present application, Figure 4The original image is shown in Figure (A). The target detection model is called to perform content recognition on the original image, and the image recognition result is obtained. Among them, the image recognition result includes 6 bounding boxes and the class label corresponding to each bounding box (for example, the class label corresponding to each bounding box is "text class"), and thus 6 text regions are obtained (that is, T = 6). As Figure 4 As shown in Figure (B), the black area is the text region.

[0153] After obtaining T text regions, the proportion of the T text regions in the original image can be calculated. Exemplarily, it is assumed that the original image includes 50,000 pixel points, and the T text regions altogether include 45,000 pixel points. Thus, the proportion of the T text regions in the original image is calculated to be 90%. Taking the first proportion threshold as 85% as an example, it can be seen that at this time, the proportion of the T text regions in the original image is greater than the first proportion threshold. Therefore, it is determined that the image type of the original image is a text image type. The text image type meets the format conversion condition. Therefore, a structure file corresponding to the original image can be generated.

[0154] It should be noted that the target detection model can be a Region Convolutional Neural Networks (R-CNN) model, or a Faster Region Convolutional Neural Networks (Faster R-CNN), or a You Only Look Once (YOLO) model, or a Single Shot MultiBox Detector (SSD), etc., which is not limited here. The first proportion threshold can also be set to 85% or other values, which is not limited here.

[0155] Secondly, in the embodiments of the present application, a method for identifying the text image type is provided. Through the above method, the target detection model can be called to quickly identify whether the original image contains text regions. For the case including text regions, the image type can also be determined according to its proportion. If it belongs to the text image type, it means that the original image contains more text content and is suitable for conversion into a corresponding structure file, thereby improving the feasibility and operability of the solution.

[0156] Optionally, on the basis of the above Figure 3 corresponding various embodiments, in another optional embodiment provided by the embodiments of the present application, generating the structure file corresponding to the original image may specifically include:

[0157] For each of the T text regions, perform optical character recognition (OCR) processing on the text region to obtain the text information corresponding to the text region;

[0158] Generate a structure file corresponding to the original image according to the text information corresponding to each of the T text regions.

[0159] In one or more embodiments, a method for generating a structure file for a text image is introduced. As can be seen from the foregoing embodiments, if the image type of the original image is a text image type, OCR processing can be further performed on each text region, thereby obtaining the text information within each text region. In addition, the position of each text region in the original image and the font attributes (such as font size, font color, font type, etc.) of the text information within each text region can also be obtained. Based on this, a structure is generated according to the text information corresponding to each of the T text regions, and a corresponding structure file is generated according to the structure.

[0160] Furthermore, in the embodiments of the present application, a method for generating a structure file for a text image is provided. Through the above method, in an IM application, users often forward the content of a screenshot text, and saving the text in the form of a picture will cause a large waste of storage space. Therefore, for pictures of the pure text content type (i.e., text image type), they can be converted into structure files for storage.

[0161] Optionally, based on the above Figure 3 In another optional embodiment provided by the embodiments of the present application on the basis of the corresponding various embodiments, the structure file includes a preset field set, and the preset field set includes a content field;

[0162] Generating a structure file corresponding to the original image according to the text information corresponding to each of the T text regions may specifically include:

[0163] Determine the text information corresponding to the content field according to the text information corresponding to each of the T text regions;

[0164] If the preset field set further includes a region position field, determine the region position information corresponding to the region position field according to the T text regions;

[0165] If the preset field set further includes a font size field, determine the font size information corresponding to the font size field according to the text information corresponding to each of the T text regions;

[0166] If the preset field set further includes a region width field, determine the region width information corresponding to the region width field according to the T text regions;

[0167] If the preset field set further includes a font color field, determine the font color information corresponding to the font color field according to the text information corresponding to each of the T text regions.

[0168] If the preset field set further includes a font bold field, determine the font bold information corresponding to the font bold field according to the text information corresponding to each of the T text regions.

[0169] If the preset field set further includes a font italic field, determine the font italic information corresponding to the font italic field according to the text information corresponding to each of the T text regions.

[0170] If the preset field set further includes a font type field, determine the font type information corresponding to the font type field according to the text information corresponding to each of the T text regions.

[0171] If the preset field set further includes a picture type field, determine the text picture type corresponding to the picture type field.

[0172] Generate a structure file corresponding to the original picture according to the text information and at least one of the region position information, font size information, region width information, font color information, font bold information, font italic information, font type information, and text picture type.

[0173] In one or more embodiments, a method for generating a structure file corresponding to a text picture type is introduced. As can be seen from the foregoing embodiments, the structure file includes a preset field set, and the preset field set includes at least a content field. That is, after OCR processing, the text information in each text region can be obtained, and the text information is the value corresponding to the content field.

[0174] Specifically, for the sake of understanding, please refer to Figure 4 , taking Figure 4 the original image shown in (A) in

[0175]

[0176]

[0177]

[0178] where "type" represents the picture type field. The value of "type" is "text", that is, it represents the text picture type.

[0179] Among them, "content" represents the content field, and the value of "content" is the text information within each text area.

[0180] Among them, "location" represents the area location field, and the value of "location" is the area location information of each text area, that is, the starting point location of the text area.

[0181] Among them, "font-size" represents the font size field, and the value of "font-size" is the font size information of the text information within each text area, that is, the font size.

[0182] Among them, "width" represents the area width field, and the value of "width" is the area width information of each text area, that is, the area width of the text area.

[0183] It should be noted that the above preset field set is only an illustration. In actual situations, the preset field set may also include other fields.

[0184] Exemplarily, a font color field "colour" can be added, and the value of "colour" is the font color information of the text information within each text area, such as black, red, gray, etc.

[0185] Exemplarily, a font bold field "bold" can be added, and the value of "bold" is the font bold information of the text information within each text area, such as bold (true), non-bold (false).

[0186] Exemplarily, a font italic field "italics" can be added, and the value of "italics" is the font italic information of the text information within each text area, such as italic (true), non-italic (false).

[0187] Exemplarily, a font type field "font-type" can be added, and the value of "font-type" is the font type information of the text information within each text area, such as Song typeface, Boldface, Regular script, etc.

[0188] Furthermore, in the embodiments of the present application, a method for generating a structure file corresponding to a text picture type is provided. Through the above method, compared with the original picture of the text picture type, the structure file only occupies a small amount of storage space, thereby greatly reducing the storage space of the text type picture.

[0189] Optionally, on the basis of the above Figure 3 corresponding various embodiments, in another optional embodiment provided by the embodiments of the present application, it may further include:

[0190] Obtain the target structure file;

[0191] If the target structure file includes a text image type, perform rendering processing according to the target structure file to obtain a first target image, where the first target image displays at least one text segment.

[0192] In one or more embodiments, a method for rendering the first target image is introduced. As can be seen from the foregoing embodiments, images of the text type are saved in the form of a structure file. When a user needs to view an image, the structure file can be read and the first target image can be rendered in real time. Among them, the display effect of the first target image is related to the preset field set included in the structure file. The effect of rendering the first target image will be described below with reference to the drawings.

[0193] Exemplarily, in one case, for ease of understanding, please refer to Figure 5 , Figure 5 , which is a schematic diagram of the original image and the first target image in an embodiment of the present application. Figure 5 In (A) of [the figure] is shown the original image. Assume that the preset field set includes a content field, an area position field, a font size field, an area width field, and an image type field. Thus, the first target image as shown in (B) of [the figure] is rendered. It can be seen that the first target image is similar to the original image in terms of the content and the content display method, that is, it includes a number of text segments. Figure 5

[0194] Exemplarily, in another case, for ease of understanding, please refer to Figure 6 , Figure 6 , which is another schematic diagram of the original image and the first target image in an embodiment of the present application. Figure 6 In (A) of [the figure] is shown the original image. Assume that the preset field set includes a content field, an area position field, a font size field, an area width field, a font type field, and an image type field. Thus, the first target image as shown in (B) of [the figure] is rendered. It can be seen that the first target image is similar to the original image in terms of the content and the content display method, that is, it includes a number of text segments. Figure 6

[0195] It can be understood that after testing, storing the original image in the form of an image occupies about 615 KB of space, while storing it in the form of a structure file occupies about 1,694 bytes, with a ratio of about 1:372, which can greatly reduce the storage space occupied by the original image. When a structure file of the text image type is recognized, a preview page is generated in real time for the user by reading the content of the structure file.

[0196] ​​Secondly, in the embodiments of the present application, a method for rendering a first target picture is provided. Through the above method, pictures of text types are saved in the form of structure files, and the first target picture can be rendered by reading the structure files. Thereby, not only can the occupation of storage space be greatly reduced, but also the loss of message content can be avoided. In addition, to a certain extent, it can also relieve the anxiety of users caused by excessive occupation of storage space when using IM applications.

[0197] Optionally, based on the above Figure 3 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, content recognition is performed on the original picture to obtain the picture type of the original picture, which may specifically include:

[0198] Performing content matching on the original picture using a first template picture to obtain a content matching result;

[0199] If the content matching result indicates a successful match, it is determined that the picture type of the original picture is a session picture type, where the session picture type meets the format conversion condition.

[0200] In one or more embodiments, a method for identifying the session picture type is introduced. As can be seen from the foregoing embodiments, for whether the picture type of the original picture is a session picture type, feature matching determination can be performed through the method of image template matching, where the feature can be a component unique to the chat interface, and this unique component is referred to as the "first template picture" here. For ease of understanding, please refer to Figure 7 , Figure 7 which is a schematic diagram of the first template picture in the embodiments of the present application. As shown in the figure, assuming that the original picture of the session picture type is a user session screenshot, exemplarily, the first template picture may include an emoji and an add symbol.

[0201] Specifically, for ease of understanding, taking the Figure 7 shown first template picture as an example, the original picture is matched by using the template matching (matchTemplat) function provided by OpenCV. Further, taking the matching method of the normalized correlation coefficient (TM_CCOEFF_NORMED) as an example, please refer to Figure 8 , Figure 8 which is a schematic diagram of implementing template matching on the original picture using the first template picture in the embodiments of the present application. Figure 8 In (A) of which, the template matching diagram is shown, where the brighter the white area, the greater the possibility of a match. By taking the matching points greater than a certain threshold (this threshold can be determined by developers based on experience, for example, 0.8) for the template matching diagram, the result as shown in Figure 8The effect diagram shown in Figure (B). If there are white bright spots, it means that there are corresponding characteristic elements. Since Figure 8 There is a white bright spot in the lower right corner of the effect diagram shown in Figure (B), so it indicates that the original picture matches the first template picture successfully. Figure 8 Figure (C) shows the original picture. Among them, the rectangular area circled in the lower right corner of the original picture is the part that matches the first template picture successfully, that is, the content matching result indicates success. Therefore, Figure 8 The original picture shown in Figure (C) belongs to the session picture type. The session picture type meets the format conversion conditions, so a structure file corresponding to the original picture can be generated.

[0202] It should be noted that in practical applications, the correlation coefficient matching method (TM_CCOEFF), or the normalized cross-correlation matching method (TM_CCORR_NORMED), or the correlation matching method (TM_CCORR), or the normalized variance matching method (TM_SQDIFF_NORMED), or the variance matching method (TM_SQDIFF) can also be used, which is not limited here.

[0203] Secondly, in the embodiments of the present application, a method for identifying the session picture type is provided. Through the above method, the template matching method can quickly identify whether the original picture contains the first template picture. If the first template picture is included, it is determined to belong to the session picture type. That is, it means that the original picture contains chat text content and is suitable for conversion into a corresponding structure file, thereby improving the feasibility and operability of the solution.

[0204] Optionally, on the basis of the above Figure 3 In another optional embodiment provided by the embodiments of the present application corresponding to each of the above embodiments, generating the structure file corresponding to the original picture may specifically include:

[0205] Perform optical character recognition (OCR) processing on the original picture to obtain M text messages, where M is an integer greater than or equal to 1;

[0206] Perform content matching on the original picture using the second template picture to obtain N session areas, where each session area includes a text message, and N is an integer greater than or equal to 1;

[0207] According to the N session areas, obtain the avatar corresponding to each session area;

[0208] Generate the structure file corresponding to the original picture according to the M text messages and the avatar corresponding to each session area.

[0209] In one or more embodiments, a method for generating a structure file for a conversation picture is introduced. If the picture type of the original picture is a conversation picture type, feature matching determination can be further performed by means of image template matching. Here, the feature can be a conversation bubble unique to the chat interface, and this unique conversation bubble is referred to as the "second template picture" herein. For ease of understanding, please refer to Figure 9 , Figure 9 which is a schematic diagram of the second template picture in an embodiment of the present application. As shown in the figure, assume that the original picture of the conversation picture type is a user conversation screenshot. Exemplarily, the second template picture can be a conversation bubble. At the same time, perform OCR processing on the original picture to obtain M text messages. Among them, the text messages within the conversation area are the chat content. Assume that N conversation areas are detected, then N of the M text messages belong to the chat content.

[0210] Specifically, for ease of understanding, please refer to Figure 10 , Figure 10 which is a schematic diagram of identifying a conversation area based on the original image in an embodiment of the present application. Figure 10 In (A) of it, the original image is shown. Based on this, taking the second template picture shown in Figure 9 as an example, use the matchTemplate function provided by OpenCV to match the original picture, and obtain 3 conversation areas (that is, N = 3) shown in (B) of Figure 10 . Among them, the black area is the conversation area. Based on this, the corresponding avatar can be obtained according to the position of the conversation area. Generate a structure based on the M text messages (that is, including the text messages corresponding to each conversation area) and the avatar corresponding to each conversation area, and generate a corresponding structure file according to the structure.

[0211] It can be understood that if it is necessary to detect whether it belongs to a voice message, feature matching determination can also be performed by means of image template matching. Here, the feature can be a voice identifier unique to the voice message, for example, an identifier of a "small speaker" or an identifier of other patterns, etc.

[0212] Again, in an embodiment of the present application, a method for generating a structure file for a conversation picture is provided. Through the above method, in an IM application, users often forward chat content by taking screenshots, and saving text in the form of pictures will cause a large waste of storage space. Therefore, for pictures of the conversation picture type, they can be converted into structure files for storage.

[0213] Optionally, in the above Figure 3Based on the corresponding respective embodiments, in another alternative embodiment provided by the embodiments of the present application, according to N session regions, obtaining the avatar corresponding to each session region may specifically include:

[0214] For each of the N session regions, obtaining the avatar corresponding to the session region according to a preset offset;

[0215] It may further include:

[0216] Performing a hash calculation on the avatar corresponding to each session region to obtain the hash information to be matched for each avatar;

[0217] If the hash information to be matched for the avatar successfully matches the avatar hash information stored locally in the terminal device, then using the hash information to be matched for the avatar as the avatar hash information of the avatar;

[0218] If the hash information to be matched for the avatar fails to match the avatar hash information stored locally in the terminal device, then storing the avatar and using the hash information to be matched for the avatar as the avatar hash information of the avatar.

[0219] In one or more embodiments, a method for storing avatars is introduced. As can be seen from the foregoing embodiments, the position of the session region can be obtained by detecting the characteristics of the session bubble. Thus, the avatar can be obtained through a fixed offset (i.e., the preset offset) between the user avatar and the session region. Here, the avatar is an avatar picture obtained after cropping.

[0220] Specifically, after obtaining the avatar corresponding to each session region, a hash calculation can be performed on each avatar to obtain the hash information to be matched for each avatar. Based on this, the hash information to be matched for each avatar is matched with the avatar hash information stored locally in the terminal device. For ease of understanding, please refer to Table 1, which is a schematic of the avatar hash information stored locally in the terminal device.

[0221] Table 1

[0222] Avatar Avatar hash information Avatar A 16er7ew41c61a8e4a1465f56we41 Avatar B 34e1s61ewf4s1a715fd4we4fwes5 Avatar C 456fsda16d6aew71swe78921fsa1 Avatar D 75dg724g87y73a18e4ba564g54q7 Avatar E 95op1g54h1nn5h4g18gf111e2134

[0223] Based on this, assuming that the hash information to be matched for a certain avatar is "16er7ew41c61a8e4a1465f56we41", at this time, the hash information to be matched for this avatar successfully matches the avatar hash information of "Avatar A". Assuming that the hash information to be matched for a certain avatar is "8yr46s4d4af6sqgdfh49a74644d4q", at this time, the hash information to be matched for this avatar fails to match the avatar hash information stored locally in the terminal device. Therefore, the avatar is stored locally in the terminal device, and the hash information to be matched for the avatar is used as the avatar hash information of the avatar and stored in the corresponding list.

[0224] It should be noted that, in actual situations, the hash value distance between the hash information to be matched and the avatar hash information can also be calculated. If the hash value distance is less than the distance threshold, it indicates that the two match successfully.

[0225] Furthermore, in the embodiments of the present application, a method for storing avatars is provided. Through the above method, the avatars in the session pictures can also be stored locally in the terminal device. Thus, when rendering with the structure file, pictures containing avatars can be rendered, achieving an effect closer to the actual screenshots.

[0226] Optionally, based on the corresponding embodiments above, Figure 3 in another optional embodiment provided by the embodiments of the present application, the structure file includes a preset field set, and the preset field set includes a message field and an avatar field;

[0227] Generating a structure file corresponding to the original picture according to the M text messages and the avatars corresponding to each session area may specifically include:

[0228] Determining the N text messages corresponding to the message field according to the M text messages;

[0229] Determining the N avatar hash information corresponding to the avatar field according to the avatars corresponding to each session area;

[0230] If the preset field set further includes a sender field, determining the N sender information corresponding to the sender field according to the positional relationship between each session area and the corresponding avatar;

[0231] If the preset field set further includes a picture type field, determining the session picture type according to the picture type field;

[0232] If the preset field set further includes an object name field, determining the object name information corresponding to the object name field according to the M text messages;

[0233] If the preset field set further includes an announcement content field, determining the announcement information corresponding to the announcement content field according to the M text messages;

[0234] If the preset field set further includes a voice message field, determining the voice type information corresponding to the voice message field according to the N text messages;

[0235] Generating a structure file corresponding to the original picture according to the N text messages and the N avatar hash information, and according to at least one of the N sender information, the session picture type, the object name information, the announcement information, and the voice type information.

[0236] In one or more embodiments, a method for generating a structure file corresponding to a session picture type is introduced. As can be seen from the foregoing embodiments, the structure file includes a preset field set, and the preset field set includes at least a message field and an avatar field. That is, after OCR processing, the text information in each session area can be obtained, and the text information is the value corresponding to the content field. After hash calculation, the avatar corresponding to each session area can be obtained, and the avatar hash information is the value corresponding to the avatar field.

[0237] Specifically, for the sake of easy understanding, please refer again to Figure 10 , taking Figure 10 the original image shown in (A) in

[0238]

[0239]

[0240] as an example. After template matching, 3 session areas are obtained. Further, after OCR recognition, the text information included in each session area is determined, and after offset calculation, the avatar corresponding to each session area is determined. Based on this, a structure can be expressed as:

[0241] Among them, "type" represents the picture type field. The value of "type" is "chat", that is, it represents the session picture type.

[0242] Among them, "title" represents the object name field. The value of "title" is the object name information, that is, the chat object name.

[0243] Among them, "bulletin" represents the announcement content field. The value of "bulletin" is the announcement information, and the announcement information belongs to one of the M text information.

[0244] Among them, "content" represents the content field, and each content field also includes multiple fields, including but not limited to sender information, avatar field, message field, and voice class information.

[0245] Among them, "avatar" represents the avatar field, and the value of "avatar" is the avatar hash information. Each chat session corresponds to one avatar hash information. Therefore, N chat sessions correspond to N avatar hash information. If the terminal device already stores the avatar corresponding to the avatar hash information locally, the avatar is directly read. Otherwise, the avatar hash information and its corresponding avatar are stored locally on the terminal device.

[0246] "msg" represents the message field, and the value of "msg" is the text information, that is, the chat content within the session area. Each chat session corresponds to one text information. Therefore, N chat sessions correspond to N text information.

[0247] It should be noted that the above preset field set is only an example. In actual situations, the preset field set may also include other fields.

[0248] Exemplarily, a voice message field "voice" can be added. The value of "voice" is voice type information. Each chat session corresponds to one voice type information. For example, it is voice type information (true) or non-voice type information (false).

[0249] Furthermore, in the embodiments of the present application, a method for generating a structure file corresponding to the session picture type is provided. Through the above method, compared with the original picture of the session picture type, the structure file only occupies a small amount of storage space, thereby greatly reducing the storage space of the session type picture.

[0250] Optionally, based on the above Figure 3 corresponding various embodiments, in another optional embodiment provided by the embodiments of the present application, it may further include:

[0251] Obtain the target structure file;

[0252] If the target structure file includes the session picture type, perform rendering processing according to the target structure file to obtain a second target picture, where the second target picture displays the avatar, the text information within the session area, and the object name information.

[0253] In one or more embodiments, a method for rendering the second target picture is introduced. As can be seen from the foregoing embodiments, the picture of the session type is saved in the form of a structure file. When the user needs to view the picture, the structure file can be read and rendered in real time to obtain the second target picture. Among them, the display effect of the second target picture is related to the preset field set included in the structure file. The effect of rendering the second target picture will be described below with reference to the drawings.

[0254] Exemplarily, in one case, for the sake of understanding, please refer to Figure 11 ,Figure 11 This is a schematic diagram of the original image and the second target image in the embodiments of the present application. Figure 11 In (A) of [Figure reference not provided], the original image is shown. Assume that the preset field set includes an object name field, an announcement content field, a sender field, an avatar field, a message field, and a picture type field. Thus, the second target image as shown in Figure 11 (B) of [Figure reference not provided] is rendered. It can be seen that the content and content display method included in the second target image are similar to those of the original image, that is, it includes an avatar, text information in the conversation area, object name information, announcement information, etc.

[0255] Exemplarily, in another case, for ease of understanding, please refer to Figure 12 , Figure 12 This is another schematic diagram of the original image and the second target image in the embodiments of the present application. Figure 12 In (A) of [Figure reference not provided], the original image is shown. Assume that the preset field set includes an object name field, an announcement content field, a sender field, an avatar field, a message field, a voice message field, and a picture type field. Thus, the second target image as shown in Figure 12 (B) of [Figure reference not provided] is rendered. It can be seen that the content and content display method included in the second target image are similar to those of the original image, that is, it includes an avatar, text information in the conversation area, object name information, announcement information, and voice type information, etc.

[0256] It can be understood that in an IM application, although users can share conversation content, due to the convenience and intuitiveness of taking screenshots, many users will take screenshots to share conversation content. Since the conversation picture contains some component color information in the conversation interface, although the text content of the communication is not much, the finally generated picture has a large storage space. After testing, the original image stored in picture form occupies about 247 KB of space, while stored in a structure file form, it occupies about 1110 bytes, with a ratio of about 1:227.86, which can greatly reduce the storage space occupied by the original image. When a structure file of the conversation picture type is recognized, a preview page is generated in real time for the user by reading the content of the structure file.

[0257] Secondly, in the embodiments of the present application, a method for rendering the second target image is provided. By the above method, pictures of the conversation type are saved in the form of structure files, and the second target image can be rendered by reading the structure files. Thus, not only can the occupation of storage space be greatly reduced, but also the loss of message content can be avoided. In addition, to a certain extent, it can also relieve the anxiety of users caused by excessive occupation of storage space when using IM applications.

[0258] Optionally, in the above Figure 3Based on the corresponding respective embodiments, in another alternative embodiment provided by the embodiments of the present application, content recognition is performed on the original picture to obtain the picture type of the original picture, which may specifically include:

[0259] Call the target detection model to perform content recognition on the original picture to obtain a picture recognition result;

[0260] If the picture recognition result indicates that the original picture includes an information code area, determine the proportion of the information code area in the original picture;

[0261] If the proportion of the information code area in the original picture is greater than or equal to the second proportion threshold, determine that the picture type of the original picture is the information code picture type, where the information code picture type meets the format conversion condition.

[0262] In one or more embodiments, a method for identifying the information code picture type is introduced. As can be seen from the foregoing embodiments, the target detection model can be directly called to perform content recognition on the original picture. Among them, the input of the target detection model is a picture, and the output is a bounding box and the class label corresponding to the bounding box. In this embodiment, the bounding box is the bounding box corresponding to the information code area, and the class label corresponding to the bounding box is "information code class".

[0263] Specifically, for ease of understanding, please refer to Figure 13 , Figure 13 is a schematic diagram for identifying the information code area based on the original image in the embodiments of the present application. Figure 13 In (A) of, the original image is shown. The target detection model is called to perform content recognition on the original picture to obtain a picture recognition result. Among them, the picture recognition result includes 1 bounding box and the class label corresponding to the bounding box (for example, the class label is "information code class (or, two-dimensional code class)"), and thus the information code area shown in (B) of Figure 13 is obtained, where the black area is the information code area.

[0264] After obtaining the information code area, the proportion of the information code area in the original picture can be calculated. Exemplarily, assume that the original picture includes 50,000 pixel points, and the information code area includes a total of 40,000 pixel points. Thus, the proportion of the information code area in the original picture is calculated to be 80%. Taking the second proportion threshold as 70% as an example, it can be seen that at this time, the proportion of the information code area in the original picture is greater than the second proportion threshold. Therefore, it is determined that the picture type of the original picture is the information code picture type. The information code picture type meets the format conversion condition. Therefore, a structure file corresponding to the original picture can be generated.

[0265] It should be noted that the target detection model can be an R-CNN model, or a Faster R-CNN, or a YOLO, or an SSD, etc., which is not limited here. The information code can include a QR code or a barcode, etc., which is not limited here. The second proportion threshold can also be set to 95% or other values, which is not limited here.

[0266] Secondly, in the embodiments of the present application, a method for identifying the type of information code picture is provided. Through the above method, by calling the target detection model, it is possible to quickly identify whether the original picture contains an information code area. For the case where an information code area is included, the picture type can be further determined according to its proportion. If it belongs to the information code picture type, it means that the original picture contains an information code and is suitable for conversion into a corresponding structure file, thereby improving the feasibility and operability of the solution.

[0267] Optionally, based on the above Figure 3 In another optional embodiment provided by the embodiments of the present application on the basis of the corresponding respective embodiments, generating the structure file corresponding to the original picture may specifically include:

[0268] Call the information code decoder to decode the information code in the information code area to obtain text information, where the text information is the link address corresponding to the information code;

[0269] Generate the structure file corresponding to the original picture according to the text information.

[0270] In one or more embodiments, a method for generating a structure file for an information code picture is introduced. As can be seen from the foregoing embodiments, if the picture type of the original picture is the information code picture type, the information code decoder can be further called to identify the link address corresponding to the information code and the information code size. Based on this, a structure is generated according to the text information (i.e., the link address) corresponding to the information code area, and a corresponding structure file is generated according to the structure.

[0271] Again, in the embodiments of the present application, a method for generating a structure file for an information code picture is provided. Through the above method, in the IM application, users often need to capture information code pictures for display, and saving the information code in the form of a picture will cause a large waste of storage space. Therefore, for pictures of the information code content type (i.e., the information code picture type), they can be converted into structure files for storage.

[0272] Optionally, based on the above Figure 3 In another optional embodiment provided by the embodiments of the present application on the basis of the corresponding respective embodiments, the structure file includes a preset field set, and the preset field set includes a content field;

[0273] Generate a structure file corresponding to the original image according to the text information, which may specifically include:

[0274] Determine the text information corresponding to the content field according to the text information;

[0275] If the preset field set further includes a picture type field, determine the information code picture type corresponding to the picture type field;

[0276] If the preset field set further includes a region width field, determine the region width information corresponding to the region width field according to the information code region;

[0277] If the preset field set further includes a region height field, determine the region height information corresponding to the region height field according to the information code region;

[0278] Generate a structure file corresponding to the original image according to the text information and at least one of the information code picture type, region width information, and region height information.

[0279] In one or more embodiments, a method for generating a structure file for an information code picture type is introduced. As can be seen from the foregoing embodiments, the structure file includes a preset field set, and the preset field set at least includes a content field. That is, the text information obtained after parsing by the information code decoder, and the text information is the value corresponding to the content field.

[0280] Specifically, for the sake of understanding, please refer to again Figure 13 , taking Figure 13 the original image shown in (A) in as an example, after target detection, the information code region is obtained, and further, the text information included in the information code region is determined by decoder parsing. Based on this, a structure can be expressed as:

[0281]

[0282] Among them, "type" represents the picture type field. The value of "type" is "qrcode", that is, it represents the information code picture type.

[0283] Among them, "content" represents the content field, and the value of "content" is the text information in the information code region, that is, the link address.

[0284] Among them, "width" represents the region width field, and the value of "width" is the region width information of the information code region, that is, the region width of the information code region.

[0285] Among them, "height" represents the region height field, and the value of "height" is the region height information of the information code region, that is, the region height of the information code region.

[0286] It should be noted that the above set of preset fields is only for illustration. In actual situations, the set of preset fields may also include other fields, which are not limited here.

[0287] Furthermore, in the embodiments of the present application, a method for generating a structure file for an information code picture type is provided. Through the above method, compared with the original picture of the information code picture type, the structure file only occupies a small amount of storage space, thus greatly reducing the storage space of the information code type picture.

[0288] Optionally, based on the corresponding respective embodiments above, in another optional embodiment provided by the embodiments of the present application, it may further include: Figure 3 Obtain a target structure file;

[0289] If the target structure file includes an information code picture type, perform rendering processing according to the target structure file to obtain a third target picture, where the third target picture displays an information code.

[0290] In one or more embodiments, a method for rendering to obtain a third target picture is introduced. As can be seen from the foregoing embodiments, the picture of the information code type is saved in the form of a structure file. When the user needs to view the picture, the structure file can be read and rendered in real time to obtain a third target picture. Among them, the display effect of the third target picture is related to the set of preset fields included in the structure file. The effect of rendering to obtain the third target picture will be described below with reference to the drawings.

[0291] Exemplarily, for ease of understanding, please refer to

[0292] For example, for ease of understanding, please refer to Figure 14 , Figure 14 which is a schematic diagram of the original picture and the third target picture in the embodiments of the present application. Figure 14 In (A) of it, the original image is shown. Assume that the set of preset fields includes a content field, a region width field, a region height field, and a picture type field. Thus, the third target picture as shown in Figure 14 in (B) of it is rendered. It can be seen that the third target picture is similar to the content and content display method included in the original picture, that is, it includes an information code (for example, a QR code).

[0293] It can be understood that after testing, the space occupied by the original picture including the QR code when generating a Portable Network Graphics (PNG) format is approximately 6,600 bytes, while the storage in the form of a structure file occupies approximately 80 bytes, with a ratio of approximately 1:82.5, which can greatly reduce the storage space occupied by the original picture. When a structure file of the information code picture type is recognized, the information code is generated in real time for the user by reading the content of the structure file.

[0294] Secondly, in the embodiments of the present application, a method for rendering a third target picture is provided. Through the above method, pictures of the information code type are saved in the form of structure files, and the third target picture can be rendered by reading the structure files. Thus, not only can the occupation of storage space be greatly reduced, but also the loss of message content can be avoided. In addition, to a certain extent, it can also relieve the anxiety of users when using IM applications due to excessive occupation of storage space.

[0295] Optionally, based on the corresponding embodiments above, in another optional embodiment provided by the embodiments of the present application, obtaining the original picture to be processed may specifically include: Figure 3 Displaying a space optimization message and a space optimization control, where the space optimization message is used to prompt the available space to be released;

[0296] Responding to a selection operation on the space optimization control to obtain the original picture to be processed;

[0297] It may also include:

[0298] Deleting the original picture from the local of the terminal device and uploading the original picture to a cloud server or other devices.

[0299]

[0300] In one or more embodiments, a method for optimizing the storage space of a device is introduced. As can be seen from the foregoing embodiments, in order to alleviate the problem of excessive storage space occupied by text-based pictures (for example, text type pictures, conversation type pictures, and information code type pictures), text-based pictures can also be optimized.

[0301] Figure 15 It can be understood that pictures usually include pure picture content, conversation screenshots, news text pictures, notification text pictures, and QR code screenshots, etc. Among them, text-based pictures can be conversation pictures stored in an IM application. Please refer to Figure 15 , Figure 16 which is a schematic diagram of an interface for storing conversation pictures in the IM application in the embodiments of the present application. As shown in the figure, the pictures shown in the figure are all pictures stored during the chat process. Text-based pictures can also be pictures stored in an album application. Please refer to Figure 16 ,Figure 16 This is a schematic diagram of the interface for storing pictures in the album application in the embodiment of the present application. As shown in the figure, the pictures shown in the figure are all pictures stored by the user.

[0302] Specifically, three ways to optimize text-based pictures will be introduced below with reference to the illustrations.

[0303] Method 1:

[0304] Exemplarily, for ease of understanding, please refer to Figure 17 , Figure 17 This is a schematic diagram of migrating the original picture in the embodiment of the present application. As shown in Figure (A) of Figure 17 , A1 is used to indicate the space optimization message, where the space optimization message prompts that the available space to be released is 3.25 Gigabytes (GB), and A2 is used to indicate the space optimization control. After the user has authorized the operation of uploading the picture to the cloud server, if the space optimization control indicated by A2 is clicked, the interface shown in Figure (B) of Figure 17 will be displayed. At this time, the original picture has been deleted from the local terminal device and uploaded to the cloud server.

[0305] Method 2:

[0306] Exemplarily, for ease of understanding, please refer to Figure 18 , Figure 18 This is another schematic diagram of migrating the original picture in the embodiment of the present application. As shown in Figure (A) of Figure 18 , B1 is used to indicate the space optimization message, where the space optimization message prompts that the available space to be released is 3.25 GB, and B2 is used to indicate the space optimization control. If the user agrees to upload the picture to the cloud server, then click the space optimization control indicated by B2. Thus, the interface shown in Figure (B) of Figure 18 will be displayed. At this time, the original picture has been deleted from the local terminal device and uploaded to the cloud server.

[0307] Method 3:

[0308] Exemplarily, for ease of understanding, please refer to Figure 19 , Figure 19 This is another schematic diagram of migrating the original picture in the embodiment of the present application. As shown in Figure (A) of Figure 19 , C1 is used to indicate the space optimization message, where the space optimization message prompts that the available space to be released is 3.25 GB, and C2 is used to indicate the space optimization control. If the user agrees to transfer the picture to another terminal device, then click the space optimization control indicated by A2. Thus, the interface shown in Figure 19The interface shown in Figure (B). At this time, the original picture has been deleted locally from the terminal device and uploaded to the cloud server.

[0309] The user can select to optimize the storage space for historical pictures. After the user opens the option, the text-based pictures will be saved in the form of a structure file. Optionally, the user can also be provided with the option to save the original pictures to other devices or the cloud server for backup storage, and only the structure file needs to be stored locally on the terminal device, and the corresponding target pictures are displayed.

[0310] Secondly, in the embodiments of the present application, a method for optimizing the storage space of the device is provided. Through the above method, not only can the storage space of the session pictures in the IM application be optimized, but also the storage space of the pictures in the local album of the terminal device can be optimized. The text-based pictures are converted into structure files composed of text attributes, and when displayed, the original information is ensured not to be lost, and at the same time, the style information of the text-based pictures is not lost. The average occupied space is only about 1 / 200 of that of picture storage, which can greatly reduce the large storage space occupied by text-based pictures. At the same time, it supports uploading the original pictures to the cloud server, which is convenient for users to view.

[0311] Optionally, on the basis of the above Figure 3 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, after storing the structure file, it may further include:

[0312] In response to a send operation for the target picture, send the target picture to the target terminal device;

[0313] Or,

[0314] After storing the structure file, it may further include:

[0315] In response to a send operation for the target picture, send the structure file to the target terminal device so that the target picture can be obtained through rendering processing according to the target structure file.

[0316] In one or more embodiments, a method for realizing data transmission based on a structure file is introduced. As can be seen from the foregoing embodiments, after the storage space is optimized, the original pictures can be deleted locally from the terminal device, and the structure files are stored locally on the terminal device. Based on this, when data is transmitted, the target picture or the structure file can be selected for transmission.

[0317] Specifically, for the sake of easy understanding, please refer to Figure 20 , Figure 20This is a schematic diagram of an interface for transmitting a target picture in an embodiment of the present application. As shown in the figure, D1 is used to indicate the original picture, and D2 is used to indicate the target picture. Exemplarily, a user can select several target pictures and transmit these target pictures to a target terminal device. Based on this, the target terminal device directly displays the target pictures. Exemplarily, a user can select several target pictures and transmit the corresponding structure files of these target pictures to the target terminal device. Based on this, the target terminal device renders the structure files, thereby displaying the corresponding target pictures.

[0318] Secondly, in an embodiment of the present application, a method for implementing data transmission based on a structure file is provided. Through the above method, a user can choose to directly send a target file or a structure file. Since both the target file and the structure file occupy less storage space. Therefore, when transmitting a target file or a structure file in an IM application, the amount of data transmission can be reduced, transmission resources can be saved, and the performance overhead of the terminal device can be reduced.

[0319] The following will describe in detail the picture processing device in the present application. Please refer to Figure 21 , Figure 21 This is a schematic diagram of an embodiment of the picture processing device in an embodiment of the present application. The picture processing device 30 includes:

[0320] An acquisition module 310, configured to acquire an original picture to be processed;

[0321] The acquisition module 310 is further configured to perform content recognition on the original picture to obtain the picture type of the original picture;

[0322] A generation module 320, configured to generate a structure file corresponding to the original picture if the picture type of the original picture meets the format conversion condition, where the structure file includes at least one text information, and the at least one text information is text content recognized based on the original picture;

[0323] A storage module 330, configured to store the structure file, where the structure file is used to generate a target picture.

[0324] In an embodiment of the present application, a picture processing device is provided. By using the above device, format conversion is performed on an original picture containing text information, that is, the original picture is recognized and relevant text information is extracted, and then a structure file is generated based on the text information. Thus, the terminal device only needs to store the structure file locally. It can be seen that the structure file mainly records text information. Therefore, compared with the original picture, the structure file occupies less storage space, thereby effectively alleviating the problem that pictures occupy too much storage space.

[0325] Optionally, in the above Figure 21Based on the corresponding embodiment, in another embodiment of the image processing apparatus 30 provided by the embodiments of the present application,

[0326] An acquisition module 310, specifically configured to call a target detection model to perform content recognition on an original image to obtain an image recognition result;

[0327] If the image recognition result indicates that the original image includes T text regions, determine the proportion of the T text regions in the original image, where T is an integer greater than or equal to 1;

[0328] If the proportion of the T text regions in the original image is greater than or equal to a first proportion threshold, determine that the image type of the original image is a text image type, where the text image type meets the format conversion condition.

[0329] In the embodiments of the present application, an image processing apparatus is provided. By using the above apparatus, the target detection model can be called to quickly identify whether the original image contains a text region. For the case where a text region is included, the image type can be further determined according to its proportion. If it belongs to the text image type, it means that the original image contains more text content and is suitable for conversion into a corresponding structure file, thereby improving the feasibility and operability of the solution.

[0330] Optionally, based on the corresponding embodiment above, in another embodiment of the image processing apparatus 30 provided by the embodiments of the present application, Figure 21 Based on the corresponding embodiment above, in another embodiment of the image processing apparatus 30 provided by the embodiments of the present application,

[0331] A generation module 320, specifically configured to perform optical character recognition (OCR) processing on each of the T text regions to obtain text information corresponding to the text region;

[0332] Generate a structure file corresponding to the original image according to the text information corresponding to each of the T text regions.

[0333] In the embodiments of the present application, an image processing apparatus is provided. In an IM application, users often forward the content of a screenshot text, and saving the text in the form of a picture will cause a large amount of storage space waste. Therefore, for a picture of the pure text content type (i.e., the text image type), it can be converted into a structure file for storage.

[0334] Optionally, based on the corresponding embodiment above, in another embodiment of the image processing apparatus 30 provided by the embodiments of the present application, the structure file includes a preset field set, and the preset field set includes a content field; Figure 21 Based on the corresponding embodiment above, in another embodiment of the image processing apparatus 30 provided by the embodiments of the present application, the structure file includes a preset field set, and the preset field set includes a content field;

[0335] A generation module 320, specifically configured to determine the text information corresponding to the content field according to the text information corresponding to each of the T text regions;

[0336] If the preset field set further includes a region position field, determine the region position information corresponding to the region position field according to the T text regions;

[0337] If the preset field set further includes a font size field, determine the font size information corresponding to the font size field according to the text information corresponding to each of the T text regions;

[0338] If the preset field set further includes a region width field, determine the region width information corresponding to the region width field according to the T text regions;

[0339] If the preset field set further includes a font color field, determine the font color information corresponding to the font color field according to the text information corresponding to each of the T text regions;

[0340] If the preset field set further includes a font bold field, determine the font bold information corresponding to the font bold field according to the text information corresponding to each of the T text regions;

[0341] If the preset field set further includes a font italic field, determine the font italic information corresponding to the font italic field according to the text information corresponding to each of the T text regions;

[0342] If the preset field set further includes a font type field, determine the font type information corresponding to the font type field according to the text information corresponding to each of the T text regions;

[0343] If the preset field set further includes a picture type field, determine the text picture type corresponding to the picture type field;

[0344] Generate a structure file corresponding to the original picture according to the text information and at least one of the region position information, font size information, region width information, font color information, font bold information, font italic information, font type information, and text picture type.

[0345] In an embodiment of the present application, a picture processing device is provided. By using the above device, compared with the original picture of the text picture type, the structure file only occupies a small amount of storage space, thereby greatly reducing the storage space of the text type picture.

[0346] Optionally, in the above Figure 21Based on the corresponding embodiment, in another embodiment of the image processing apparatus 30 provided in the embodiments of the present application, the image processing apparatus 30 further includes a processing module 340;

[0347] The obtaining module 310 is further configured to obtain a target structure file;

[0348] The processing module 340 is configured to, if the target structure file includes a text image type, perform rendering processing according to the target structure file to obtain a first target image, where the first target image displays at least one text segment.

[0349] In the embodiments of the present application, an image processing apparatus is provided. By using the above apparatus, images of text types are saved in the form of structure files, and the first target image can be rendered by reading the structure files. Thus, not only can the occupation of storage space be greatly reduced, but also the loss of message content can be avoided. In addition, to a certain extent, it can also relieve the anxiety of users caused by excessive occupation of storage space when using IM applications.

[0350] Optionally, based on the corresponding embodiment above, in another embodiment of the image processing apparatus 30 provided in the embodiments of the present application, Figure 21 The obtaining module 310 is specifically configured to perform content matching on the original image by using a first template image to obtain a content matching result;

[0351] If the content matching result indicates a successful match, it is determined that the image type of the original image is a session image type, where the session image type meets the format conversion condition.

[0352] In the embodiments of the present application, an image processing apparatus is provided. By using the above apparatus, the template matching method can quickly identify whether the original image contains the first template image. If the first template image is included, it is determined to belong to the session image type. That is, it means that the original image contains chat text content and is suitable for conversion into a corresponding structure file, thereby improving the feasibility and operability of the solution.

[0353] Optionally, based on the corresponding embodiment above, in another embodiment of the image processing apparatus 30 provided in the embodiments of the present application,

[0354] The generating module 320 is specifically configured to perform optical character recognition (OCR) processing on the original image to obtain M text messages, where M is an integer greater than or equal to 1; Figure 21 Based on the corresponding embodiment above, in another embodiment of the image processing apparatus 30 provided in the embodiments of the present application,

[0355]

[0356] ​Perform content matching on the original image using a second template image to obtain N conversation regions, where each conversation region includes a text message, and N is an integer greater than or equal to 1;

[0357] Based on the N conversation regions, obtain the avatar corresponding to each conversation region;

[0358] Generate a structure file corresponding to the original image according to the M text messages and the avatar corresponding to each conversation region.

[0359] In the embodiments of the present application, an image processing device is provided. Using the above device, in an IM application, users often forward chat content by taking screenshots, and saving text in the form of images will cause a large waste of storage space. Therefore, for images of the conversation picture type, they can be converted into structure files for storage.

[0360] Optionally, based on the corresponding embodiments above, in another embodiment of the image processing device 30 provided by the embodiments of the present application, Figure 21 The generation module 320 is specifically configured to, for each of the N conversation regions, obtain the avatar corresponding to the conversation region according to a preset offset;

[0361] The generation module 320 is specifically configured to, for each of the N conversation regions, obtain the avatar corresponding to the conversation region according to a preset offset;

[0362] The processing module 340 is further configured to perform a hash calculation on the avatar corresponding to each conversation region to obtain the hash information to be matched for each avatar;

[0363] The processing module 340 is further configured to, if the hash information to be matched for the avatar matches the avatar hash information stored locally in the terminal device, use the hash information to be matched for the avatar as the avatar hash information;

[0364] The storage module 330 is further configured to, if the hash information to be matched for the avatar fails to match the avatar hash information stored locally in the terminal device, store the avatar and use the hash information to be matched for the avatar as the avatar hash information.

[0365] In the embodiments of the present application, an image processing device is provided. Using the above device, the avatars in the conversation pictures can also be stored locally in the terminal device. Thus, when rendering using the structure file, pictures containing avatars can be rendered, achieving an effect closer to the actual screenshots.

[0366] Optionally, based on the corresponding embodiments above, in another embodiment of the image processing device 30 provided by the embodiments of the present application, the structure file includes a preset field set, and the preset field set includes a message field and an avatar field; Figure 21 The storage module 330 is further configured to, if the hash information to be matched for the avatar fails to match the avatar hash information stored locally in the terminal device, store the avatar and use the hash information to be matched for the avatar as the avatar hash information.

[0367] A generation module 320, specifically configured to determine N pieces of text information corresponding to message fields according to M pieces of text information;

[0368] According to the avatars corresponding to each session area, determine N avatar hash information corresponding to the avatar field;

[0369] If the preset field set further includes a sender field, then according to the positional relationship between each session area and the corresponding avatar, determine N sender information corresponding to the sender field;

[0370] If the preset field set further includes a picture type field, then determine the session picture type according to the picture type field;

[0371] If the preset field set further includes an object name field, then according to M pieces of text information, determine the object name information corresponding to the object name field;

[0372] If the preset field set further includes an announcement content field, then according to M pieces of text information, determine the announcement information corresponding to the announcement content field;

[0373] If the preset field set further includes a voice message field, then according to N pieces of text information, determine the voice type information corresponding to the voice message field;

[0374] Generate a structure file corresponding to the original picture according to the N pieces of text information and the N avatar hash information, and according to at least one of the N sender information, the session picture type, the object name information, the announcement information, and the voice type information.

[0375] In an embodiment of the present application, a picture processing device is provided. By using the above device, compared with the original picture of the session picture type, the structure file only occupies a small amount of storage space, thereby greatly reducing the storage space of the session type picture.

[0376] Optionally, on the basis of the above Figure 21 corresponding embodiment, in another embodiment of the picture processing device 30 provided in the embodiment of the present application,

[0377] An acquisition module 310 is further configured to acquire a target structure file;

[0378] A processing module 340 is further configured to, if the target structure file includes a session picture type, perform rendering processing according to the target structure file to obtain a second target picture, where the second target picture displays an avatar, text information within the session area, and object name information.

[0379] In an embodiment of the present application, an image processing device is provided. By using the above device, pictures of the session type are saved in the form of a structure file, and a second target picture can be rendered by reading the structure file. Thus, not only can the occupation of storage space be greatly reduced, but also the loss of message content can be avoided. In addition, to a certain extent, it can also relieve the anxiety of users caused by excessive occupation of storage space when using IM applications.

[0380] Optionally, based on the corresponding embodiment above, in another embodiment of the image processing device 30 provided in the embodiment of the present application, Figure 21 the acquisition module 310 is specifically configured to call a target detection model to perform content recognition on the original picture to obtain a picture recognition result;

[0381] If the picture recognition result indicates that the original picture includes an information code area, determine the proportion of the information code area in the original picture;

[0382] If the proportion of the information code area in the original picture is greater than or equal to a second proportion threshold, determine that the picture type of the original picture is an information code picture type, where the information code picture type meets the format conversion condition.

[0383]

[0384] In an embodiment of the present application, an image processing device is provided. By using the above device, the target detection model can be called to quickly identify whether the original picture contains an information code area. For the case where the information code area is included, the picture type can also be determined according to its proportion. If it belongs to the information code picture type, it means that the original picture contains an information code and is suitable for conversion into a corresponding structure file, thereby improving the feasibility and operability of the solution.

[0385] Optionally, based on the corresponding embodiment above, in another embodiment of the image processing device 30 provided in the embodiment of the present application, Figure 21 the generation module 320 is specifically configured to call an information code decoder to decode the information code in the information code area to obtain text information, where the text information is the link address corresponding to the information code;

[0386] Generate a structure file corresponding to the original picture according to the text information.

[0387]

[0388] In an embodiment of the present application, an image processing device is provided. By using the above device, in an IM application, users often need to intercept information code pictures for display, and saving the information code in the form of a picture will cause a large amount of waste of storage space. Therefore, for pictures of the information code content type (i.e., the information code picture type), they can be converted into structure files for storage.

[0389] Optionally, based on the above Figure 21 In another embodiment of the image processing apparatus 30 provided in the embodiments of the present application, on the basis of the corresponding embodiment, the structure file includes a preset field set, and the preset field set includes a content field;

[0390] The generating module 320 is specifically configured to determine the text information corresponding to the content field according to the text information;

[0391] If the preset field set further includes a picture type field, then determine the information code picture type corresponding to the picture type field;

[0392] If the preset field set further includes a region width field, then determine the region width information corresponding to the region width field according to the information code region;

[0393] If the preset field set further includes a region height field, then determine the region height information corresponding to the region height field according to the information code region;

[0394] Generate a structure file corresponding to the original picture according to the text information and at least one of the information code picture type, region width information, and region height information.

[0395] In the embodiments of the present application, an image processing apparatus is provided. By using the above apparatus, compared with the original picture of the information code picture type, the structure file only occupies a small amount of storage space, thereby greatly reducing the storage space of the information code type picture.

[0396] Optionally, based on the above Figure 21 In another embodiment of the image processing apparatus 30 provided in the embodiments of the present application, on the basis of the corresponding embodiment,

[0397] The obtaining module 310 is further configured to obtain a target structure file;

[0398] The processing module 340 is further configured to, if the target structure file includes an information code picture type, perform rendering processing according to the target structure file to obtain a third target picture, where the third target picture displays an information code.

[0399] In the embodiments of the present application, an image processing apparatus is provided. By using the above apparatus, the picture of the information code type is saved in the form of a structure file, and the third target picture can be rendered by reading the structure file. Thus, not only can the occupation of storage space be greatly reduced, but also the loss of message content can be avoided. In addition, to a certain extent, it can also relieve the anxiety of users caused by excessive occupation of storage space when using IM applications.

[0400] Optionally, based on the above Figure 21Based on the corresponding embodiment, in another embodiment of the image processing apparatus 30 provided in the embodiments of the present application,

[0401] An obtaining module 310 is specifically configured to display a space optimization message and a space optimization control, where the space optimization message is used to prompt the available space that can be released;

[0402] In response to a selection operation on the space optimization control, obtain an original image to be processed;

[0403] A processing module 340 is further configured to delete the original image locally from the terminal device and upload the original image to a cloud server or other devices.

[0404] In the embodiments of the present application, an image processing apparatus is provided. By using the above apparatus, not only can the storage space of session images in the IM application be optimized, but also the storage space of images in the local album of the terminal device can be optimized. The text-based image is converted into a structure file composed of text attributes. When displayed, the original information can be ensured not to be lost, and at the same time, the style information of the text-based image can be ensured not to be lost. The average occupied space is only about 1 / 200 of that of the image storage, which can greatly reduce the large storage space occupied by the text-based images. At the same time, it supports uploading the original image to the cloud server, facilitating the user to view.

[0405] Optionally, based on the above Figure 21 Based on the corresponding embodiment, in another embodiment of the image processing apparatus 30 provided in the embodiments of the present application,

[0406] After storing the structure file, the processing module 340 is further configured to, in response to a sending operation on the target image, send the target image to the target terminal device;

[0407] Or,

[0408] After storing the structure file, the processing module 340 is further configured to, in response to a sending operation on the target image, send the structure file to the target terminal device, so that the target image is obtained through rendering processing according to the target structure file.

[0409] In the embodiments of the present application, an image processing apparatus is provided. By using the above apparatus, the user can choose to directly send the target file or the structure file. Since both the target file and the structure file occupy less storage space. Therefore, when transmitting the target file or the structure file in the IM application, the data transmission volume can be reduced, the transmission resources can be saved, and the performance overhead of the terminal device can be reduced.

[0410] The embodiments of the present application further provide an image processing apparatus deployed on a terminal device. As Figure 22As shown, for ease of explanation, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present application. In the embodiments of the present application, a mobile phone is taken as an example of the terminal device for illustration:

[0411] Figure 22 The figure shows a block diagram of a part of the structure of a mobile phone related to the terminal device provided by the embodiments of the present application. Refer to Figure 22 , the mobile phone includes: a radio frequency (RF) circuit 410, a memory 420, an input unit 430, a display unit 440, a sensor 450, an audio circuit 460, a wireless fidelity (WiFi) module 470, a processor 480, and a power supply 490 and other components. Those skilled in the art can understand that Figure 22 the structure of the mobile phone shown in

[0412] does not limit the mobile phone, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. Figure 22 The following specifically introduces each component of the mobile phone:

[0413] The RF circuit 410 can be used to receive and send signals during information reception or call processes. Specifically, after receiving the downlink information from the base station, it is given to the processor 480 for processing; in addition, the designed uplink data is sent to the base station. Generally, the RF circuit 410 includes but is not limited to antennas, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 410 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0414] The memory 420 can be used to store software programs and modules. The processor 480 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 420. The memory 420 may mainly include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.); the data storage area can store the data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 420 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0415] The input unit 430 can be used to receive input digital or character information and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 430 may include a touch panel 431 and other input devices 432. The touch panel 431, also known as a touch screen, can collect the touch operations of the user on or near it (such as the operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 431), and drive the corresponding connection device according to a pre-set program. Optionally, the touch panel 431 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 480, and can receive and execute the commands sent by the processor 480. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 431. In addition to the touch panel 431, the input unit 430 may also include other input devices 432. Specifically, the other input devices 432 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a mouse, a joystick, etc.

[0416] The display unit 440 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 440 may include a display panel 441. Optionally, the display panel 441 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 431 can cover the display panel 441. When the touch panel 431 detects a touch operation on or near it, it is transmitted to the processor 480 to determine the type of touch event. Subsequently, the processor 480 provides a corresponding visual output on the display panel 441 according to the type of touch event. Although in Figure 22 the touch panel 431 and the display panel 441 are implemented as two independent components to realize the input and input functions of the mobile phone, in some embodiments, the touch panel 431 and the display panel 441 can be integrated to realize the input and output functions of the mobile phone.

[0417] The mobile phone may further include at least one sensor 450, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 441 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 441 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that the mobile phone can also be configured with, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be elaborated here.

[0418] The audio circuit 460, the speaker 461, and the microphone 462 can provide an audio interface between the user and the mobile phone. The audio circuit 460 can transmit the electrical signal converted from the received audio data to the speaker 461, and the speaker 461 converts it into a sound signal for output; on the other hand, the microphone 462 converts the collected sound signal into an electrical signal, which is received by the audio circuit 460 and converted into audio data. After the audio data is output and processed by the processor 480, it is sent through the RF circuit 410 to, for example, another mobile phone, or the audio data is output to the memory 420 for further processing.

[0419] WiFi belongs to short-range wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the WiFi module 470. It provides users with wireless broadband Internet access. Although Figure 22The WiFi module 470 is shown, but it can be understood that it does not belong to the essential components of the mobile phone and can be completely omitted within the scope of not changing the essence of the invention as needed.

[0420] The processor 480 is the control center of the mobile phone, connecting various parts of the entire mobile phone through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 420, and by calling data stored in the memory 420, it executes various functions of the mobile phone and processes data. Optionally, the processor 480 may include one or more processing units; optionally, the processor 480 may integrate an application processor and a modem processor, where the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 480 either.

[0421] The mobile phone also includes a power supply 490 (such as a battery) for supplying power to each component. Optionally, the power supply can be logically connected to the processor 480 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system.

[0422] Although not shown, the mobile phone may also include a camera, a Bluetooth module, etc., which will not be elaborated here.

[0423] The steps performed by the terminal device in the above embodiments can be based on the Figure 22 shown terminal device structure.

[0424] In the embodiments of the present application, a computer device is also provided, including a memory and a processor. When the processor executes a computer program stored in the memory, the steps of the methods described in the foregoing embodiments are implemented.

[0425] In the embodiments of the present application, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the methods described in the foregoing embodiments are implemented.

[0426] In the embodiments of the present application, a computer program product is also provided, including a computer program. When the computer program is executed by a processor, the steps of the methods described in the foregoing embodiments are implemented.

[0427] It can be understood that in the specific implementation manners of the present application, data related to user information, conversation records, original pictures, etc. are involved. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.

[0428] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0429] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0430] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0431] In addition, each functional unit in various embodiments of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0432] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store computer programs.

[0433] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for image processing, characterized in that, Including: Obtain the original image to be processed; Perform content recognition on the original image to obtain the image type of the original image; If the image type of the original image meets the format conversion condition, generate a structure file corresponding to the original image, where the structure file includes at least one text information, and the at least one text information is the text content recognized based on the original image; If the image type is a text image type, the structure file further includes: at least one of area position information, font size information, area width information, font color information, font bold information, font italic information, font type information, and text image type; If the image type is a conversation image type, the structure file further includes: avatar hash information, and at least one of sender information, conversation image type, object name information, announcement information, and voice type information; If the image type is an information code image type, the structure file further includes: at least one of information code image type, area width information, and area height information; Store the structure file, where the structure file is used to generate a target image, and the content display mode of the target image is similar to that of the original image.

2. The method according to claim 1, wherein The performing content recognition on the original image to obtain the image type of the original image includes: Call a target detection model to perform content recognition on the original image to obtain an image recognition result; If the image recognition result indicates that the original image includes T text areas, determine the proportion of the T text areas in the original image, where T is an integer greater than or equal to 1; If the proportion of the T text areas in the original image is greater than or equal to a first proportion threshold, determine that the image type of the original image is a text image type, where the text image type meets the format conversion condition.

3. The method according to claim 2, wherein The generating the structure file corresponding to the original image includes: For each of the T text areas, perform optical character recognition (OCR) processing on the text area to obtain the text information corresponding to the text area; Generate the structure file corresponding to the original image according to the text information corresponding to each of the T text areas.

4. The method according to claim 3, wherein The structure file includes a preset field set, and the preset field set includes a content field; The generating the structure file corresponding to the original image according to the text information corresponding to each of the T text areas includes: Determine the text information corresponding to the content field according to the text information corresponding to each of the T text areas; If the preset field set further includes an area position field, determine the area position information corresponding to the area position field according to the T text areas; If the preset field set further includes a font size field, determine the font size information corresponding to the font size field according to the text information corresponding to each of the T text areas; If the preset field set further includes a region width field, determine the region width information corresponding to the region width field according to the T text regions; If the preset field set further includes a font color field, determine the font color information corresponding to the font color field according to the text information corresponding to each of the T text regions; If the preset field set further includes a font bold field, determine the font bold information corresponding to the font bold field according to the text information corresponding to each of the T text regions; If the preset field set further includes a font italic field, determine the font italic information corresponding to the font italic field according to the text information corresponding to each of the T text regions; If the preset field set further includes a font type field, determine the font type information corresponding to the font type field according to the text information corresponding to each of the T text regions; If the preset field set further includes a picture type field, determine the text picture type corresponding to the picture type field; Generate the structure file corresponding to the original picture according to the text information and at least one of the region position information, the font size information, the region width information, the font color information, the font bold information, the font italic information, the font type information, and the text picture type.

5. The method according to claim 1, characterized in that, The method further includes: Obtain a target structure file; If the target structure file includes a text picture type, perform rendering processing according to the target structure file to obtain a first target picture, where at least one text segment is displayed in the first target picture.

6. The method according to claim 1, wherein The content recognition of the original picture to obtain the picture type of the original picture includes: Perform content matching on the original picture with a first template picture to obtain a content matching result; If the content matching result indicates successful matching, determine that the picture type of the original picture is a session picture type, where the session picture type meets the format conversion condition.

7. The method according to claim 6, characterized in that, The generating the structure file corresponding to the original picture includes: Perform optical character recognition (OCR) processing on the original picture to obtain M text information, where M is an integer greater than or equal to 1; Perform content matching on the original picture with a second template picture to obtain N session regions, where each session region includes a text information, and N is an integer greater than or equal to 1; According to the N session regions, obtain the avatar corresponding to each session region; Generate the structure file corresponding to the original picture according to the M text information and the avatar corresponding to each session region.

8. The method according to claim 7, characterized in that, The obtaining the avatar corresponding to each session region according to the N session regions includes: For each of the N session regions, obtain the avatar corresponding to the session region according to a preset offset; The method further includes: Perform hash calculation on the avatars corresponding to each session area to obtain the hash information to be matched for each avatar; If the hash information to be matched for the avatar successfully matches the avatar hash information stored locally in the terminal device, then use the hash information to be matched for the avatar as the avatar hash information of the avatar; If the hash information to be matched for the avatar fails to match the avatar hash information stored locally in the terminal device, then store the avatar and use the hash information to be matched for the avatar as the avatar hash information of the avatar.

9. The method according to claim 7 or 8, characterized in that, The structure file includes a preset field set, and the preset field set includes a message field and an avatar field; The generating the structure file corresponding to the original picture according to the M text messages and the avatars corresponding to each session area includes: Determine the N text messages corresponding to the message field according to the M text messages; Determine the N avatar hash information corresponding to the avatar field according to the avatars corresponding to each session area; If the preset field set further includes a sender field, then determine the N sender information corresponding to the sender field according to the positional relationship between each session area and the corresponding avatar; If the preset field set further includes a picture type field, then determine the session picture type according to the picture type field; If the preset field set further includes an object name field, then determine the object name information corresponding to the object name field according to the M text messages; If the preset field set further includes an announcement content field, then determine the announcement information corresponding to the announcement content field according to the M text messages; If the preset field set further includes a voice message field, then determine the voice type information corresponding to the voice message field according to the N text messages; Generate the structure file corresponding to the original picture according to the N text messages and the N avatar hash information, and according to at least one of the N sender information, the session picture type, the object name information, the announcement information and the voice type information.

10. The method according to claim 1, wherein The method further includes: Obtain a target structure file; If the target structure file includes a session picture type, then perform rendering processing according to the target structure file to obtain a second target picture, where the second target picture displays avatars, text messages in the session area and object name information.

11. The method according to claim 1, wherein The performing content recognition on the original picture to obtain the picture type of the original picture includes: Call a target detection model to perform content recognition on the original picture to obtain a picture recognition result; If the picture recognition result indicates that the original picture includes an information code area, then determine the proportion of the information code area in the original picture; If the proportion of the information code area in the original picture is greater than or equal to a second proportion threshold, then determine that the picture type of the original picture is an information code picture type, where the information code picture type meets the format conversion condition.

12. The method according to claim 11, wherein The generating the structure file corresponding to the original picture includes: Call the information code decoder to decode the information code in the information code area to obtain text information, where the text information is the link address corresponding to the information code; Generate the structure file corresponding to the original picture according to the text information.

13. The method according to claim 12, wherein The structure file includes a preset field set, and the preset field set includes a content field; The generating the structure file corresponding to the original picture according to the text information includes: Determine the text information corresponding to the content field according to the text information; If the preset field set further includes a picture type field, determine the information code picture type corresponding to the picture type field; If the preset field set further includes a region width field, determine the region width information corresponding to the region width field according to the information code area; If the preset field set further includes a region height field, determine the region height information corresponding to the region height field according to the information code area; Generate the structure file corresponding to the original picture according to the text information and at least one of the information code picture type, the region width information, and the region height information.

14. The method according to claim 1, wherein The method further includes: Obtain a target structure file; If the target structure file includes an information code picture type, perform rendering processing according to the target structure file to obtain a third target picture, where the information code is displayed on the third target picture.

15. The method according to claim 1, wherein The obtaining the original picture to be processed includes: Display a space optimization message and a space optimization control, where the space optimization message is used to prompt the available space to be released; Respond to the selection operation on the space optimization control to obtain the original picture to be processed; The method further includes: Delete the original picture from the local of the terminal device and upload the original picture to the cloud server or other devices.

16. The method according to claim 1, characterized in that, After storing the structure file, the method further includes: Respond to the sending operation on the target picture and send the target picture to the target terminal device; Or, After storing the structure file, the method further includes: Respond to the sending operation on the target picture and send the structure file to the target terminal device, so that the target picture is obtained by performing rendering processing according to the structure file.

17. An image processing device, characterized in that, Includes: An obtaining module, configured to obtain an original picture to be processed; The obtaining module is further configured to perform content recognition on the original picture to obtain the picture type of the original picture; A generation module, configured to generate a structure file corresponding to the original image if the image type of the original image meets the format conversion condition, where the structure file includes at least one text information, and the at least one text information is the text content recognized based on the original image; if the image type is a text image type, the structure file further includes at least one of region position information, font size information, region width information, font color information, font bold information, font italic information, font type information, and text image type; if the image type is a conversation image type, the structure file further includes at least one of avatar hash information, sender information, conversation image type, object name information, announcement information, and voice type information; if the image type is an information code image type, the structure file further includes at least one of information code image type, region width information, and region height information. A storage module, configured to store the structure file, where the structure file is used to generate a target image, and the content display mode of the target image is similar to that of the original image.

18. The device according to claim 17, wherein The obtaining module is specifically configured to: Call a target detection model to perform content recognition on the original image to obtain an image recognition result; If the image recognition result indicates that the original image includes T text regions, determine the proportion of the T text regions in the original image, where T is an integer greater than or equal to 1; If the proportion of the T text regions in the original image is greater than or equal to a first proportion threshold, determine that the image type of the original image is a text image type, where the text image type meets the format conversion condition.

19. The device according to claim 18, wherein The generation module is specifically configured to: For each of the T text regions, perform optical character recognition (OCR) processing on the text region to obtain the text information corresponding to the text region; Generate the structure file corresponding to the original image according to the text information corresponding to each of the T text regions.

20. The device according to claim 19, wherein The structure file includes a preset field set, and the preset field set includes a content field; the generation module is specifically configured to: Determine the text information corresponding to the content field according to the text information corresponding to each of the T text regions; If the preset field set further includes a region position field, determine the region position information corresponding to the region position field according to the T text regions; If the preset field set further includes a font size field, determine the font size information corresponding to the font size field according to the text information corresponding to each of the T text regions; If the preset field set further includes a region width field, determine the region width information corresponding to the region width field according to the T text regions; If the preset field set further includes a font color field, determine the font color information corresponding to the font color field according to the text information corresponding to each of the T text regions; If the preset field set further includes a font bold field, determine the font bold information corresponding to the font bold field according to the text information corresponding to each of the T text regions; If the preset field set further includes a font italic field, determine the font italic information corresponding to the font italic field according to the text information corresponding to each of the T text regions; If the preset field set further includes a font type field, determine the font type information corresponding to the font type field according to the text information corresponding to each of the T text regions; If the preset field set further includes a picture type field, determine the text picture type corresponding to the picture type field; Generate the structure file corresponding to the original picture according to the text information and at least one of the region position information, the font size information, the region width information, the font color information, the font bold information, the font italic information, the font type information, and the text picture type.

21. The device according to claim 17, wherein The device further includes: a processing module; The obtaining module is further configured to obtain a target structure file; The processing module is configured to, if the target structure file includes a text picture type, perform rendering processing according to the target structure file to obtain a first target picture, where at least one text segment is displayed in the first target picture.

22. The device according to claim 17, characterized in that, The obtaining module is specifically configured to: Perform content matching on the original picture using a first template picture to obtain a content matching result; If the content matching result indicates a successful match, determine that the picture type of the original picture is a session picture type, where the session picture type meets the format conversion condition.

23. The device according to claim 22, wherein, The generating module is specifically configured to: Perform optical character recognition (OCR) processing on the original picture to obtain M pieces of text information, where M is an integer greater than or equal to 1; Perform content matching on the original picture using a second template picture to obtain N session regions, where each session region includes a piece of text information, and N is an integer greater than or equal to 1; Obtain the avatar corresponding to each of the N session regions according to the N session regions; Generate the structure file corresponding to the original picture according to the M pieces of text information and the avatar corresponding to each session region.

24. The device according to claim 23, characterized in that, The device further includes a processing module; The generating module is specifically configured to: for each of the N session regions, obtain the avatar corresponding to the session region according to a preset offset; The processing module is configured to perform a hash calculation on the avatar corresponding to each session region to obtain the hash information to be matched for each avatar; The processing module is configured to, if the hash information to be matched for the avatar matches the avatar hash information locally stored in the terminal device, use the hash information to be matched for the avatar as the avatar hash information of the avatar; The storage module is further configured to store the avatar and use the to-be-matched hash information of the avatar as the avatar hash information of the avatar if the to-be-matched hash information of the avatar fails to match the avatar hash information stored locally in the terminal device.

25. The device according to claim 23 or 24, characterized in that, The structure file includes a preset field set, and the preset field set includes a message field and an avatar field; the generating module is specifically configured to: Determine the N text messages corresponding to the message field according to the M text messages; Determine the N avatar hash information corresponding to the avatar field according to the avatars corresponding to each session area; If the preset field set further includes a sender field, determine the N sender information corresponding to the sender field according to the positional relationship between each session area and the corresponding avatar; If the preset field set further includes a picture type field, determine the session picture type according to the picture type field; If the preset field set further includes an object name field, determine the object name information corresponding to the object name field according to the M text messages; If the preset field set further includes a notice content field, determine the notice information corresponding to the notice content field according to the M text messages; If the preset field set further includes a voice message field, determine the voice type information corresponding to the voice message field according to the N text messages; Generate the structure file corresponding to the original picture according to the N text messages and the N avatar hash information, and according to at least one of the N sender information, the session picture type, the object name information, the notice information, and the voice type information.

26. The device according to claim 17, wherein The device further includes: a processing module; The obtaining module is further configured to obtain a target structure file; The processing module is configured to perform a rendering process on the target structure file to obtain a second target picture if the target structure file includes a session picture type, where the second target picture displays an avatar, text messages in the session area, and object name information.

27. The device according to claim 17, characterized in that, The obtaining module is specifically configured to: Call a target detection model to perform content recognition on the original picture to obtain a picture recognition result; If the picture recognition result indicates that the original picture includes an information code area, determine the proportion of the information code area in the original picture; If the proportion of the information code area in the original picture is greater than or equal to a second proportion threshold, determine that the picture type of the original picture is an information code picture type, where the information code picture type meets the format conversion condition.

28. The device according to claim 27, wherein The generating module is specifically configured to: Call an information code decoder to decode the information code in the information code area to obtain text information, where the text information is the link address corresponding to the information code; Generate the structure file corresponding to the original picture according to the text information.

29. The device according to claim 28, characterized in that, The structure file includes a preset field set, and the preset field set includes a content field; the generating module is specifically configured to: Determine the text information corresponding to the content field according to the said text information; If the preset field set further includes a picture type field, determine the information code picture type corresponding to the picture type field; If the preset field set further includes a region width field, determine the region width information corresponding to the region width field according to the information code region; If the preset field set further includes a region height field, determine the region height information corresponding to the region height field according to the information code region; Generate the structure file corresponding to the original picture according to the said text information and at least one of the information code picture type, the region width information, and the region height information; 30. The device according to claim 17, characterized in that, The device further includes: a processing module; The obtaining module is further configured to obtain a target structure file; The processing module is configured to, if the target structure file includes an information code picture type, perform rendering processing according to the target structure file to obtain a third target picture, where the information code is displayed on the third target picture; 31. The device according to claim 17, characterized in that, The device further includes: a processing module; The obtaining module is specifically configured to: Display a space optimization message and a space optimization control, where the space optimization message is used to prompt the available space to be released; In response to a selection operation on the space optimization control, obtain the original picture to be processed; The processing module is configured to delete the original picture locally from the terminal device and upload the original picture to a cloud server or other device; 32. The device according to claim 17, wherein The device further includes: a processing module; The processing module is configured to, after storing the structure file, in response to a sending operation on the target picture, send the target picture to a target terminal device; Or, The processing module is configured to, after storing the structure file, in response to a sending operation on the target picture, send the structure file to a target terminal device, so that the target picture is obtained by performing rendering processing according to the structure file; 33. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 16 are implemented; 34. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 16 are implemented; 35. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 16 are implemented;

Citation Information

Patent Citations

  • Image processing method and device, computer equipment and storage medium

    CN111652272A

  • Information conversion method and device and storage medium

    CN113791860A