Image processing method, device, equipment and computer readable storage medium
By structurally deconstructing and semantically expanding text images and automatically matching them with a material library, the problem of low efficiency and poor quality in generating special effects text images has been solved, achieving efficient and rich generation of special effects text images.
Patent Information
- Application Number
- CN202110655189.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-11
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2041-11-01
AI Technical Summary
In existing technologies, the generation efficiency of special effects text images is low, the efficiency of manual design by designers is low, the design solutions are limited, the performance effect is poor, and the design specifications of different designers are difficult to unify, resulting in low production efficiency.
By deconstructing the text structure of the text image to be processed, identifying key structures, obtaining and expanding semantic information, and automatically matching it using a preset material library, special effects text images are generated.
It improves the efficiency and performance of special effects text and images, saves on manual selection and material screening costs, and enables richer design solutions.
Smart Images

Figure CN113821663B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to artificial intelligence technology, and in particular to an image processing method and device, equipment and a computer readable storage medium. BACKGROUND
[0002] At present, information flow products of the Internet, such as public number text and image, commodity advertisements, etc., often enrich their product performance through a large number of advertisement images, information flow head images, etc., such as Figure 1 and Figure 2 the special effect text image shown in the figure and the picture material. This kind of special effect text image usually needs to be designed and produced manually by a designer, which is low in efficiency; and the design scheme of the special effect text provided by the designer is also limited, therefore, the current generation method of the special effect text image is low in efficiency, the form of the generated special effect text image is not rich enough, and the performance effect is poor. SUMMARY
[0003] The embodiments of the present application provide an image processing method, device, equipment and computer readable storage medium, which can improve the efficiency of generating a special effect text image and improve the performance effect of the special effect text image.
[0004] The technical scheme of the embodiments of the present application is as follows:
[0005] The embodiments of the present application provide an image processing method, which comprises:
[0006] Based on a preset structure category, the obtained to-be-processed text image is subjected to text structure disassembly to obtain at least one text structure;
[0007] According to a preset screening rule, the at least one text structure is screened to obtain a landmark structure;
[0008] The semantic information in the to-be-processed text image is obtained and subjected to semantic expansion to obtain expanded semantic information;
[0009] Based on the landmark structure and the expanded semantic information, the structure information and content label corresponding to each material in at least one material in a preset material library are matched to obtain a special effect material;
[0010] The landmark structure in the to-be-processed text image is covered with the special effect material to obtain a special effect text image.
[0011] The embodiments of the present application provide an image processing device, which comprises:
[0012] The structure identification module is configured to disassemble the obtained to-be-processed text image based on a preset structure category to obtain at least one text structure;
[0013] The screening module is configured to screen the at least one character structure according to a preset screening rule to obtain a landmark structure.
[0014] The semantic module is configured to obtain semantic information in the to-be-processed character image and perform semantic extension to obtain extended semantic information.
[0015] The matching module is configured to match the landmark structure and the extended semantic information with structure information and content labels corresponding to each material in at least one material in a preset material library to obtain a special effect material.
[0016] The covering module is configured to cover the landmark structure in the to-be-processed character image with the special effect material to obtain a special effect character image.
[0017] In the device, each rule in the preset screening rule corresponds to a preset rule priority, and the preset screening rule includes at least one of the following:
[0018] The landmark structure occupies an area ratio of the to-be-processed character image, which is greater than or equal to a preset area ratio threshold;
[0019] The number of landmark structures is less than or equal to a first preset number threshold;
[0020] The landmark structure is a preset structure in the preset structure category;
[0021] The landmark structure is located in a preset area in the to-be-processed character image.
[0022] In the device, the image processing device further includes an adjusting module, which is configured to adjust the preset screening rule to obtain an adjusted screening rule when the number of landmark structures is less than a second preset number threshold; the second preset number threshold is less than the first preset number threshold; the adjusted screening rule is used for dynamic relaxation processing of the preset screening rule to increase the number of landmark structures obtained according to the adjusted screening rule; the at least one character structure is re-screened according to the adjusted screening rule to obtain the landmark structure; wherein the dynamic relaxation processing includes at least one of reducing the number of used rules and reducing the preset area ratio threshold in the preset screening rule.
[0023] In the device, the semantic module is further configured to perform character content recognition on the to-be-processed character image to obtain a character sequence corresponding to the to-be-processed character image; perform word segmentation processing on the character sequence to obtain at least one word; perform word meaning extension on the at least one word to obtain at least one extended word corresponding to each word; and take each word and the at least one extended word corresponding to the word as the extended semantic information.
[0024] In the device, the semantic module is further configured to calculate, for each of the at least one word, a similarity between the each of the at least one word and each preset word vector in a preset word vector library; and determine, as an extended word corresponding to the each of the at least one word, a preset word vector whose similarity is greater than or equal to a preset similarity threshold, to obtain the at least one extended word.
[0025] In the device, the matching module is further configured to determine, from the at least one material, a set of to-be-matched materials that match the landmark structure; calculate, for each to-be-matched material in the set of to-be-matched materials, a structure score corresponding to the each to-be-matched material according to a preset structure matching weight and a prediction probability in structure information corresponding to the each to-be-matched material; calculate a similarity between a content label of the each to-be-matched material and the extended semantic information to obtain a semantic score corresponding to the each to-be-matched material in combination with a preset semantic matching weight; and filter the special effect material from the set of to-be-matched materials based on the structure score and the semantic score.
[0026] In the device, the content label is at least one identified content, and each identified content includes a label confidence. The matching module is further configured to determine, for each to-be-matched material, a candidate content label whose label confidence is greater than or equal to a preset confidence threshold from the at least one content label; and calculate a similarity between the candidate content label and the extended semantic information to obtain the semantic score corresponding to the each to-be-matched material.
[0027] In the device, the matching module is further configured to calculate, for each material in the preset material library, a structure score corresponding to the each material according to a prediction probability in structure information of the each material corresponding to the landmark structure in combination with a preset structure matching weight; calculate a similarity between a content label of the each material and the extended semantic information to obtain a semantic score corresponding to the each material in combination with a preset semantic matching weight; and filter the special effect material from the preset material library based on the structure score and the semantic score.
[0028] In the device, the image processing device further comprises a material processing module, and the material processing module is configured to: before obtaining the special effect material by matching the preset structure category with the structural information and the content label corresponding to each material in at least one material in a preset material library based on the landmark structure and the extended semantic information, perform classification prediction on the original material according to the preset structure category, to obtain the at least one prediction probability of the at least one structure corresponding to the original material; when a maximum prediction probability in the at least one prediction probability is greater than a preset probability threshold, perform image content recognition on the original material to obtain the content label corresponding to the original material; and store the at least one prediction probability of the at least one structure corresponding to the original material as the structural information, and store the original material, the structural information corresponding to the original material, and the content label in the preset material library.
[0029] In the device, the material processing module is further configured to: obtain material color information of the original material and at least one of the adaptive regions of the original material; and store the material color information and the at least one of the adaptive regions, the original material, the structural information corresponding to the original material, and the content label in the preset material library.
[0030] In the device, the matching module is further configured to: filter a candidate material set from the preset material library according to the structure score and the semantic score; obtain color information corresponding to the landmark structure or background color information of the to-be-processed text image as to-be-matched color information; and filter, from the candidate material set, a candidate material that matches at least one of regions occupied by the landmark structure, as the special effect material, according to at least one of adaptive regions and material color information of each candidate material.
[0031] In the device, the matching module is further configured to: perform summation or average calculation on the structure score and the semantic score to obtain a comprehensive score corresponding to each material in the preset material library; and filter the special effect material from the preset material library based on the comprehensive score.
[0032] In the device, the material processing module is further configured to: for each structure in the preset structure category, obtain a preset number of original materials corresponding to the structure from an original material library to obtain a sample material set; perform model training on an initial multi-classification neural network by using the sample material set to obtain a structure classification model; and perform classification on remaining original materials in the original material library by using the structure classification model to obtain the at least one prediction probability of the at least one structure corresponding to the remaining original materials, so as to complete the classification of the original material.
[0033] In the apparatus, the structure identification module is further configured to perform target detection on the to-be-processed character image based on the preset structure category, and predict a plurality of image regions corresponding to the preset structure category; each image region contains at least one region confidence; the at least one region confidence represents at least one probability that each image region corresponds to at least one structure in the preset structure category; and a structure corresponding to a region confidence greater than or equal to a preset structure confidence threshold is taken as the at least one character structure in the to-be-processed character image.
[0034] In the apparatus, each image region contains position information; the overlay module is further configured to perform image preprocessing on the special effect material to obtain preprocessed special effect material; and the preprocessed special effect material is overlaid on an image region occupied by the landmark structure according to the position information to obtain the special effect character image; the image preprocessing includes at least one of scaling processing and rotation processing; and the scaling processing is used to adjust the size of the special effect material according to the image region occupied by the landmark structure.
[0035] In the apparatus, the matching module is further configured to match landmark structures and extended semantic information with structure information and content labels corresponding to each material in at least one material in a preset material library to obtain a plurality of alternative materials; and any alternative material in the plurality of alternative materials is taken as the special effect material.
[0036] The image processing apparatus further includes a replacement module, which is configured to, after the landmark structure in the to-be-processed character image is overlaid with the special effect material to obtain a special effect character image, generate a new special effect character image according to the remaining materials in the plurality of alternative materials when receiving a special effect character image replacement instruction.
[0037] An electronic device is provided in an embodiment of the present application, and the electronic device includes:
[0038] A memory is configured to store executable instructions.
[0039] A processor is configured to execute the executable instructions stored in the memory to implement the image processing method provided in the embodiments of the present application.
[0040] A computer readable storage medium is provided in an embodiment of the present application, and the computer readable storage medium stores executable instructions, which are used to cause a processor to execute the image processing method provided in the embodiments of the present application.
[0041] The embodiments of the present application have the following beneficial effects:
[0042] The embodiment of the present application realizes automatic identification of representative structures such as Chinese strokes from the to-be-processed character image by performing character structure disassembly on the to-be-processed character image and automatically screening out a landmark structure from at least one character structure obtained by disassembly, saves the cost of selecting representative strokes from the to-be-processed character image by manual work, improves the processing efficiency of character structure selection, and further improves the efficiency of generating special effect character images. Moreover, the present application can make more abundant materials matched according to the expanded semantic information, and improve the performance effect of the special effect character image. Further, the embodiment of the present application realizes automatic matching of the landmark structure and the expanded semantic information in the preset material library, saves the cost of screening a large amount of materials by manual work, improves the matching efficiency, further improves the generation efficiency of the special effect character image, and can expand the material matching range to the entire material library, so that better special effect materials can be selected in a larger matching range, and the performance effect of the special effect character image is further improved. BRIEF DESCRIPTION OF DRAWINGS
[0043] Figure 1 is an optional effect schematic diagram of a special effect character image provided by the embodiment of the present application;
[0044] Figure 2 is an optional effect schematic diagram of a special effect character image provided by the embodiment of the present application;
[0045] Figure 3 is an optional structure schematic diagram of an image processing system architecture provided by the embodiment of the present application;
[0046] Figure 4 is an optional structure schematic diagram of an image processing device provided by the embodiment of the present application;
[0047] Figure 5 is an optional flow schematic diagram of an image processing method provided by the embodiment of the present application;
[0048] Figure 6 is an effect schematic diagram of part of a character structure provided by the embodiment of the present application;
[0049] Figure 7 is an optional flow schematic diagram of an image processing method provided by the embodiment of the present application;
[0050] Figure 8 is a process schematic diagram of classification and prediction of original materials by a structure classification model provided by the embodiment of the present application;
[0051] Figure 9 is a material schematic diagram provided by the embodiment of the present application;
[0052] Figure 10 is an optional flowchart of an image processing method provided by an embodiment of the present application;
[0053] Figure 11 is an optional flowchart of an image processing method provided by an embodiment of the present application;
[0054] Figure 12 is an exemplary application process diagram of an image processing method provided by an embodiment of the present application in an actual application scenario. DETAILED DESCRIPTION
[0055] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be described in further detail below with reference to the accompanying drawings, and the described embodiments should not be regarded as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present application.
[0056] In the following description, "some embodiments" are described, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0057] In the following description, the terms "first\second\third" are only to distinguish similar objects, and do not represent a specific order of the objects. It can be understood that "first\second\third" can be interchanged in a specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0059] The related data collection and processing in the embodiments of the present application should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of authorization of laws and regulations and the personal information subject.
[0060] Before the embodiments of the present application are described in further detail, the terms and terms involved in the embodiments of the present application are explained, and the terms and terms involved in the embodiments of the present application are applicable to the following explanations.
[0061] 1) Artificial Intelligence (AI) is the theory, method, technology and application system of using digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is the design principle and implementation method of various intelligent machines, so that the machine has the functions of perception, reasoning and decision-making.
[0062] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technology generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, automatic driving, intelligent transportation, etc.
[0063] 2) Computer Vision (CV) Computer vision is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes to identify, track and measure targets, and further process graphics so that the computer processing becomes more suitable for human eye observation or image transmission to instrument detection. As a scientific discipline, computer vision researches related theories and technologies, trying to establish artificial intelligence systems that can obtain information from images or multidimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, intelligent transportation, etc. It also includes common face recognition, fingerprint recognition and other biometric identification technologies.
[0064] 3) Machine Learning (ML) is a multi-disciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, etc. It is a subject that studies how computers simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent, and its applications are widespread in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and example-based learning technologies.
[0065] 4) Design graph source file: the source file designed by using design software such as Photoshop or Sketch, which contains the information of elements and their positions, sizes, etc. of each layer.
[0066] 5) Canvas: refers to the visible area of the design graph, and the size of the canvas is also the size of the picture.
[0067] 6) Layer / element: the design graph is composed of a plurality of elements stacked in order. Each layer is a layer. Elements refer to text, graphics, pictures, and even tables, etc. distributed in each layer.
[0068] With the research and progress of artificial intelligence technology, artificial intelligence technology is researched and applied in many fields, such as common smart home, smart wearable device, virtual assistant, smart speaker, smart marketing, unmanned driving, autonomous driving, unmanned aerial vehicle, robot, smart medical treatment, smart customer service, Internet of vehicles, autonomous driving, intelligent transportation, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0069] The scheme provided by the embodiments of the present application relates to image processing technology in computer vision of artificial intelligence, which is specifically explained by the following embodiments.
[0070] At present, the design of special effect text images by artificial labor will produce high implementation cost, and the pictures generated by manpower are limited, so it is difficult for designers to extensively involve in the massive materials in the material library, thereby leading to the design scheme of the special effect text image produced is relatively single, not enough rich and diverse, and the performance effect is poor. Moreover, the design specifications of different designers are difficult to unify, and it is difficult to ensure that the design quality meets the specifications, which needs to be reviewed and controlled in the later stage, further increasing the implementation cost and reducing the production efficiency.
[0071] The embodiments of the present application provide an image processing method, device, equipment and computer readable storage medium, which can improve the efficiency of generating special effect text images and improve the performance effect of special effect text images. The following describes an exemplary application of the electronic device provided by the embodiments of the present application. The electronic device provided by the embodiments of the present application can be implemented as a notebook computer, a tablet computer, a desktop computer, a set-top box, a mobile device (for example, a mobile phone, a portable music player, a personal digital assistant, a dedicated message device, a portable game device) and various types of user terminals. It can also be implemented as a server. The following will illustrate an exemplary application when the electronic device is implemented as a server.
[0072] Referring to Figure 3 , Figure 3is an optional architecture schematic diagram of the image processing system 100 provided by the embodiment of the present application, for realizing a special effect text image generation application, the terminal 400 (exemplarily shows the terminal 400-1 and the terminal 400-2) connects the server 200 through the network 300, and the network 300 can be a wide area network or a local area network, or a combination of the two.
[0073] The terminal 400 is used for receiving the to-be-processed text image input by the user through a function entrance, and sending the to-be-processed text image to the server 200 through the network 300. The graphical interface 410 (exemplarily shows the graphical interface 410-1 and the graphical interface 410-2) displays the function entrance of the special effect text image generation application, such as the function menu of an APP or the entrance of a web application. The server 200 is used for performing text structure disassembly on the obtained to-be-processed text image based on a preset structure category, obtaining at least one text structure; performing screening on the at least one text structure according to a preset screening rule, obtaining a landmark structure; obtaining semantic information in the to-be-processed text image and performing semantic expansion, obtaining expanded semantic information; in the preset material library of the database 500, matching the landmark structure and the expanded semantic information with structure information and content labels corresponding to each material in the at least one material, obtaining a special effect material; covering the landmark structure in the to-be-processed text image with the special effect material, obtaining a special effect text image. The server 200 sends the special effect text image to the terminal 400, and presents the special effect text image to the user through the graphical interface 410, that is, the effect of special effect beautification on the to-be-processed text image.
[0074] In some embodiments, the server 200 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDNs, and basic cloud computing services such as big data and artificial intelligence platforms. The terminal 400 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, and the like, but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in the embodiment of the present application.
[0075] Referring to Figure 4 , Figure 4 is a structure schematic diagram of the server 200 provided by the embodiment of the present application, Figure 4The illustrated server 200 includes at least one processor 210, memory 250, at least one network interface 220, and a user interface 230. The various components of server 200 are coupled together by a bus system 240, which is configured to implement any of a variety of bus architectures. By way of example, the bus system 240 can be implemented using a PCI Express bus, a HyperTransport bus, a bus architecture supported by AMD® processors, or any other suitable bus technology. The bus system 240 is used for communicating information between the various hardware components in the server 200. The bus system 240 is implemented using one or more busses, which are communicatively coupled together using various bridges, controllers, and / or adapters. Figure 4 In this illustrative example, the bus system 240 includes a data bus, a control bus, and a state bus. However, in other implementations, the bus system 240 can include any other suitable type of bus structure.
[0076] The processor 210 can be an integrated circuit chip, such as a general purpose processor, a Digital Signal Processor (DSP), or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any other suitable processing component. The processor 210 can be a microprocessor, or any other suitable processing component.
[0077] The user interface 230 includes one or more output devices 231 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 230 also includes one or more input devices 232 that facilitate user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, other input buttons and controls.
[0078] The memory 250 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, and the like. The memory 250 optionally includes one or more storage devices remotely located from the processor 210.
[0079] The memory 250 includes volatile memory or nonvolatile memory, or both. Nonvolatile memory can be read only memory (ROM), programmable ROM (PROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, or the like. Volatile memory can include random access memory (RAM), dynamic RAM (DRAM), static RAM (SRAM), fast page mode DRAM (FPM DRAM), extended data output RAM (EDO RAM), extended data output dual data rate RAM (EDO DDR RAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous DRAM (SSDRAM), video random access memory (VRAM), cache memory (including various levels), or the like. The memory 250 of the subject embodiments is intended to include any suitable type of memory.
[0080] In some embodiments, the memory 250 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or superset thereof, which are described in the examples below.
[0081] The operating system 251 includes system programs for processing various basic system services and for performing hardware dependent tasks, such as a framework layer, a core library layer, a driver layer, and the like, for implementing various basic services and processing hardware-dependent tasks;
[0082] a network communication module 252 for reaching other computing devices via one or more (wired or wireless) network interfaces 220, exemplary network interfaces 220 including: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), etc.;
[0083] a presentation module 253 for enabling presentation of information via one or more output devices 231 (e.g., display screens, speakers, etc.) associated with the user interface 230 (e.g., a user interface for operating the peripheral device and displaying content and information);
[0084] an input processing module 254 for detecting and translating one or more user inputs or interactions from one or more input devices 232.
[0085] In some embodiments, the apparatus provided by the embodiments of the present application can be implemented in software, Figure 4 An image processing apparatus 255 stored in the memory 250 is shown, which can be software in the form of programs and plug-ins, etc., including the following software modules: a structure identification module 2551, a screening module 2552, a semantic module 2553, a matching module 2554, and a covering module 2555. These modules are logical, and thus can be combined or further split according to the implemented functions.
[0086] The functions of the various modules will be described below.
[0087] In some other embodiments, the apparatus provided by the embodiments of the present application can be implemented in hardware, as an example, the apparatus provided by the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the image processing method provided by the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can use one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic elements.
[0088] The image processing method provided by the embodiments of the present application will be described in conjunction with exemplary applications and implementations of the server provided by the embodiments of the present application.
[0089] Referring toFigure 5 , Figure 5 is an optional flowchart of the image processing method provided by the embodiment of the present application, which will be described in combination with the steps shown in the figure. Figure 5
[0090] S101, based on a preset structure category, performing text structure disassembly on the obtained to-be-processed text image, to obtain at least one text structure.
[0091] The image processing method provided by the embodiment of the present application can be applied to the Internet, content information, e-commerce, enterprise business, finance and insurance, education and training, life service, and general entertainment industry scenes, and can be used in the scene of creating content such as posters, advertising pictures, self-media pictures, and promotional videos.
[0092] In the embodiment of the present application, the image processing device can directly obtain the to-be-processed text image input by the user, or generate a corresponding to-be-processed text image according to the font, font size, text content, background picture and other parameter information input by the user. The to-be-processed text image is an image file containing text.
[0093] In the embodiment of the present application, in order to realize the detection of at least one text structure from the to-be-processed text image, the image processing device can collect a sample image set containing the preset structure category according to each structure in the preset structure category, and iteratively train an initial target detection neural network model according to the sample image set to obtain a stroke recognition model. Further, the image processing device can use the pre-trained stroke recognition model to perform target detection on the to-be-processed text image based on the preset structure category, and predict a plurality of image regions corresponding to the preset structure category. Each image region contains at least one region confidence, and each region confidence represents at least one probability that each image region corresponds to at least one structure in the preset structure category. The image processing device can use the region confidence information of each image region to identify the image region with a region confidence greater than or equal to a preset structure confidence threshold as a recognized text structure, thereby obtaining at least one text structure.
[0094] In some embodiments, the stroke recognition model can be a convolutional neural network (CNN) model, which can be a fast region CNN (Fast-RCNN) model, or other neural network models, which are selected according to actual conditions, and the embodiment of the present application is not limited.
[0095] In the embodiments of the present application, the image processing apparatus can also utilize image segmentation algorithms and image matching algorithms in image processing to perform text structure disassembly on the obtained to-be-processed text image, and obtain at least one text structure. The specific selection is made according to actual conditions, and the embodiments of the present application are not limited.
[0096] In the embodiments of the present application, the preset structure category can include strokes of Chinese characters such as Chinese characters, can include letter outline shapes of English, and can also include text structures of other languages. Alternatively, the preset structure category of Chinese can also include radicals or components; the preset structure category can also be self-defined according to common use, and the specific selection is made according to actual conditions, and the embodiments of the present application are not limited.
[0097] In some embodiments, as shown in Figure 6 Figure 6 Part of the text structure obtained by the image processing apparatus performing text structure disassembly on the to-be-processed text image containing “Lichun” is shown, as shown in 61-64 in Figure 6
[0098] S102, according to the preset filtering rule, filtering at least one text structure to obtain a landmark structure.
[0099] In the embodiments of the present application, the image processing apparatus can select representative strokes as the landmark structure from the at least one text structure according to the preset filtering rule, perform special effect beautification based on the landmark structure, and generate a special effect text image.
[0100] In some embodiments, the preset filtering rule can include at least one of the following:
[0101] The area proportion of the landmark structure in the to-be-processed text image is greater than or equal to a preset area proportion threshold; the number of landmark structures is less than or equal to a first preset number threshold; the landmark structure is a preset structure in the preset structure category; and the landmark structure is located in a preset region in the to-be-processed text image.
[0102] In some embodiments, the preset area proportion threshold can be 30%, that is, whether the area proportion of the stroke in the to-be-processed text image is greater than or equal to 30% is taken as a rule. Other proportion threshold values can also be set, and the specific selection is made according to actual conditions, and the embodiments of the present application are not limited.
[0103] In some embodiments, the first preset quantity threshold can be a threshold for screening the quantity of the logo structure in a single character, such as when the first preset quantity threshold is 1, in each character contained in the to-be-processed character image, at most one stroke is selected as the logo structure; the first preset quantity threshold can also be a threshold for screening the quantity of the logo structure in multiple characters. For example, for a to-be-processed image containing multiple characters, when the first preset quantity threshold is 2, the logo structure representing at most two characters is selected. The first preset quantity threshold can also be set to other numerical values, which are selected according to actual conditions, and the embodiments of the present application are not limited.
[0104] In some embodiments, the preset structure can be any designated structure in the preset structure category, such as a square stroke structure as a preset structure in the preset screening rule. The specific selection is based on the actual situation, and the embodiments of the present application are not limited.
[0105] In some embodiments, the preset region can be a region where the logo structure is located in the to-be-processed character image, such as the left side of the character in the to-be-processed character image as the preset region. The specific selection can be based on the actual situation, and the embodiments of the present application are not limited.
[0106] In the embodiments of the present application, the image processing apparatus screens at least one character structure by one of the above-mentioned preset screening rules; or different weights or rule priorities can be preset for each rule in the preset screening rule, and at least one character structure is screened according to the priority; or the image processing apparatus can select any several rules in the preset screening rule according to the actual situation, and screen at least one character structure according to the combined screening rule, and the specific selection is based on the actual situation, and the embodiments of the present application are not limited.
[0107] It should be noted that in some embodiments, when the quantity of the logo structure obtained according to the preset screening rule is insufficient, such as when the quantity of the logo structure is less than the preset second quantity threshold, the image processing apparatus can adjust the preset screening rule, such as appropriately relaxing the preset screening rule to obtain an adjusted screening rule. Here, the second preset quantity threshold is less than the first preset quantity threshold. Further, the image processing apparatus can re-screen at least one character structure according to the adjusted screening rule to obtain the logo structure.
[0108] In some embodiments, the screening rule is adjusted for dynamic relaxation of the preset screening rule to increase the number of landmark structures obtained according to the adjusted screening rule, wherein the dynamic relaxation can include at least one of reducing the number of rules used and reducing the preset area ratio threshold in the preset screening rule. For example, for reducing the number of rules used, the image processing apparatus can randomly subtract one rule used, or subtract a rule with a lower priority according to a preset priority corresponding to each rule. The specific relaxation strategy can be selected according to actual conditions, and the embodiments of the present application are not limited.
[0109] In some embodiments, the process of the above dynamic relaxation can be an automatic and continuous dynamic adjustment process according to the number of landmark structures screened according to the adjusted screening rule. For example, when the number of landmark structures screened according to the adjusted screening rule is still less than the second preset number threshold, the image processing apparatus can further relax on the basis of the adjusted screening rule to obtain a new adjusted screening rule, and re-screen at least one text structure according to the new adjusted screening rule until a suitable landmark structure is screened.
[0110] Here, through the dynamic relaxation process, the flexibility of screening landmark text structures can be improved.
[0111] S103, obtaining semantic information in the to-be-processed text image and performing semantic extension to obtain extended semantic information.
[0112] In the embodiments of the present application, the landmark structure belongs to the shape dimension, and the image processing apparatus can obtain semantic information of the text content from the to-be-processed text image and perform semantic extension to obtain extended semantic information in the semantic content dimension.
[0113] In some embodiments, the image processing apparatus can perform text content recognition on the to-be-processed text image to obtain a character sequence corresponding to the to-be-processed text image; the image processing apparatus performs word segmentation processing on the character sequence to obtain at least one word; the image processing apparatus performs word meaning extension on the at least one word to obtain at least one extended word corresponding to each word; and the image processing apparatus takes each word and its corresponding at least one extended word as the extended semantic information.
[0114] Exemplarily, the image processing apparatus can use a general python library such as jieba to perform word segmentation on the character sequence recognized from the to-be-processed character image by using the inverse document frequency technology, for example, divide “gu yu shi jie” into two words “gu yu” and “shi jie”. The image processing apparatus can perform word sense extension on the word “gu yu” by using the word vector technology, and extend related words such as “rainwater”, “crop”, “fruit”, “garlic” and the like as at least one extended word. Alternatively, for the word “li chun”, related words such as “magpie” and “peach blossom” can be extended as at least one extended word. Alternatively, for the word “fruit”, a series of related words such as “watermelon” and “apple” can be extended as at least one extended word.
[0115] In some embodiments, the process of performing word sense extension on the at least one word by the image processing apparatus to obtain at least one extended word corresponding to each word can be implemented by the following process:
[0116] The image processing apparatus calculates the similarity between each word and each preset word vector in the preset word vector library for each word in the at least one word; and takes the preset word vector with a similarity greater than or equal to a preset similarity threshold as an extended word corresponding to each word, to obtain at least one extended word.
[0117] In some embodiments, the preset word vector library can be the Tencent_AILab_ChineseEmbedding library open sourced by AI Lab, or other word vector libraries. The method of calculating the similarity between each word and each preset word vector in the preset word vector library by the image processing apparatus can be a cosine similarity algorithm, or can use a natural language processing framework such as Gensim and Annoy to load the word vector library and output the similarity between each word and each preset word vector, or can be other similarity algorithms, which are selected according to actual conditions, and the embodiments of the present application are not limited.
[0118] In some embodiments, the data content of the word vector library can be as shown in Table 1.
[0119]
[0120] Table 1
[0121] In Table 1, for the input word, the word vector library can contain multiple preset word vectors related to the input word, i.e., similar words.
[0122] It should be noted that the process corresponding to S103 and S101-S102 is a parallel relationship, which can be executed before or after S101-S102, or simultaneously with S101-S102, and the execution order of the above two processes is not limited by the embodiments of the present application.
[0123] S104, matching the structural information and the content label corresponding to each of the at least one material in the preset material library based on the landmark structure and the extended semantic information, to obtain the special effect material.
[0124] In the embodiments of the present application, the landmark structure represents information of shape dimension, and the extended semantic information represents information of semantic content dimension. The image processing device can match in the at least one material in the preset material library according to the landmark structure and the extended semantic information, and select the material matching in both the shape dimension and the semantic content dimension as the special effect material.
[0125] In the embodiments of the present application, each material in the preset material library includes structural information and a content label. The structural information includes at least one prediction probability of each material corresponding to at least one structure in a preset structure category, and the content label includes material content information corresponding to the material. The image processing device can compare the landmark structure and the extended semantic information with the structural information and the content label of each material, match in the preset material library, and obtain the special effect material.
[0126] In the embodiments of the present application, the image processing device can match in the preset material library according to various different preset matching strategies. For example, the image processing device can preset that the weight of matching according to the extended semantic information is greater than the weight of matching according to the landmark structure, so that the special effect material obtained by matching is not inharmonious with the text content. Alternatively, the image processing device can also set different priorities for matching according to the extended semantic information and matching according to the landmark structure, so as to realize preliminary screening according to the extended semantic information and matching according to the landmark structure, or preliminary screening according to the landmark structure and matching according to the extended semantic information. Alternatively, the image processing device can also calculate a semantic matching score of each material with respect to the extended semantic information and a structural matching score of each material with respect to the landmark structure based on similarity, and obtain a comprehensive score by a series of operation and processing based on the semantic matching score and the structural matching score, and screen the special effect material according to the comprehensive score. Alternatively, the image processing device can also screen the material with the same color as the overall color of the text image to be processed or more suitable for display in the area of the landmark structure based on more dimensions such as color dimension or position dimension. Alternatively, the image processing device can set different matching thresholds according to actual conditions to screen, and the like. The specific matching strategy can be flexibly set and selected according to actual conditions, and the embodiments of the present application are not limited.
[0127] In some embodiments, the image processing device can also combine the above-mentioned various strategies to match according to the combined strategy. The specific matching strategy can be flexibly set and selected according to actual conditions, and the embodiments of the present application are not limited.
[0128] S105, cover the landmark structure in the to-be-processed text image using the special effect material to obtain a special effect text image.
[0129] In the embodiment of the application, the image processing apparatus can cover the special effect material in the form of a picture at the position corresponding to the landmark architecture in the to-be-processed text image, replace the landmark structure in the text structure with the special effect material, and obtain a special effect text image capable of presenting a combination of text and image.
[0130] In some embodiments, the to-be-processed text image can be a design graph source file containing multiple layers, the landmark structure can be an element in a certain layer such as layer 2, and the image processing apparatus can obtain the position information of each image region output by the target detection in the process of text structure disassembly by target detection in S101, and then obtain the position information corresponding to the landmark structure. In some embodiments, the position information corresponding to the landmark structure can be the coordinates of each vertex of the landmark structure. The image processing apparatus places the special effect material at the corresponding position in the upper layer of layer 2 according to the coordinates of each vertex of the landmark structure, so as to cover the landmark structure of the to-be-processed image with the special effect material in the canvas and obtain a special effect text image.
[0131] In some embodiments, the image processing apparatus will first preprocess the special effect material before covering the landmark structure with the special effect material, to obtain a preprocessed special effect material, so as to perfect the presentation effect of the special effect text image after covering. The image processing apparatus can cover the preprocessed special effect material in the image region occupied by the landmark structure according to the position information, to obtain a special effect text image; wherein the image preprocessing includes at least one of scaling processing and rotation processing.
[0132] In the embodiment of the application, the scaling processing is used to adjust the size of the special effect material according to the to-be-processed image region. For example, the image processing apparatus can adjust each vertex of the special effect material to a position beyond the vertex of the landmark structure, to completely cover the landmark structure. The rotation processing is used to adjust the angle and direction of the special effect material. For example, the image processing apparatus can determine the matching feature points between the outline of the landmark structure and the outline of the special effect material by a feature point matching method, align the special effect material with the landmark structure according to the matching feature points, and thus adjust the placement angle of the special effect material, so as to make the special effect material more coordinated with the overall text.
[0133] It can be understood that the embodiment of the present application realizes automatic identification of representative structures such as Chinese character strokes from the to-be-processed character image by performing character structure disassembly on the to-be-processed character image and automatically screening out the landmark structure from at least one character structure obtained by disassembly, saves the cost of selecting representative strokes from the to-be-processed character image by manual work, improves the processing efficiency of character structure selection, and further improves the efficiency of generating special effect character images. Moreover, the present application can make the expanded semantic information match more abundant materials, and improve the performance effect of the special effect character image. Further, the embodiment of the present application realizes automatic matching of the landmark structure and the expanded semantic information in the preset material library, saves the cost of screening a large amount of materials by manual work, improves the matching efficiency, further improves the generation efficiency of the special effect character image, and can expand the material matching range to the entire material library, so as to select a special effect material with better performance effect in a larger matching range, and further improve the performance effect of the special effect character image.
[0134] In some embodiments, the image processing apparatus can match the landmark structure and the expanded semantic information with the structure information and the content label corresponding to each of the at least one material in the preset material library to obtain a plurality of alternative materials, so that the special effect character images of a plurality of different design schemes can be generated according to the plurality of alternative materials. The image processing apparatus can store the plurality of alternative materials in a standby material pool. For the current special effect character image processing process, any one material in the standby material pool is used as a special effect material to generate a special effect character image. When the user is not satisfied with the current generated special effect character image, the user can issue a special effect character image replacement instruction through a preset human-computer interaction interface. The image processing apparatus can generate a new special effect character image according to the remaining materials in the plurality of alternative materials when the special effect character image replacement instruction is received.
[0135] It can be understood that the image processing method provided by the embodiment of the present application can be deployed in an image editing online tool or in a copy image production platform to quickly generate a beautified special effect character image according to a to-be-processed character image submitted by a user to the platform or the tool, greatly improving the efficiency of image editing and processing. Moreover, if the user is not satisfied with the design scheme of the current generated special effect character image, the user can also replace the design scheme again by using the method provided by the embodiment of the present application, improving the flexibility of special effect character image generation and enriching the performance effect of the special effect character image.
[0136] In some embodiments, based on Figure 5 , before S104, the image processing apparatus can also obtain a preset material library by performing S001-S003 as shown in Figure 7 , which will be described in combination with each step.
[0137] S001. Classify and predict the original materials according to the preset structure categories to obtain at least one prediction probability of at least one structure corresponding to the original materials.
[0138] In the embodiments of the present application, when a preset material library is newly built according to a large number of original materials in the original material library, or a small number of original materials need to be stored in the existing preset material library, the image processing device can classify and predict each original material to be processed according to the preset structure categories to obtain at least one prediction probability of at least one structure corresponding to each original material.
[0139] In some embodiments, the image processing device can use a pre-trained multi-class neural network model to implement the classification and prediction of the original materials. The image processing device can obtain a sample material set, and each sample material is marked with a corresponding annotation result belonging to a certain structure of the preset structure category. Exemplarily, when the preset structure category is the Chinese character stroke category, the annotation result of a circular material can be the structure of "mouth" or the structure of "field", and the annotation result of a material inclined from left to right can be the structure of "dot". The image processing device can use the sample material set to train the initial multi-class neural network, and iteratively adjust the network parameters of the initial multi-class neural network according to the error between the prediction result of the initial multi-class neural network for the sample materials and the annotation result in each round of training until the classification recognition accuracy of the initial multi-class neural network reaches the preset training target, and then obtain the structure classification model.
[0140] In the embodiments of the present application, the image processing device can obtain the sample material set from an external source of the original material library, or use the original materials in the original material library. For each structure in the preset structure category, obtain the original materials corresponding to each structure with a preset sample quantity from the original material library to obtain the material sample set. The image processing device can use the material sample set to train the initial multi-class neural network to obtain the structure classification model; and use the trained structure classification model to classify the remaining original materials in the original material library to obtain at least one prediction probability of at least one structure corresponding to the remaining original materials, so as to complete the classification of the original materials.
[0141] In some embodiments, the process of classifying and predicting the original materials through the structure classification model can be as Figure 8As shown, the image processing device can acquire multiple sample materials corresponding to each structure in a preset structural category, obtaining a sample material set. This sample material set is then used to pre-train an initial CNN classification network, resulting in the CNN classification network. The CNN classification network can contain convolutional layers, pooling layers, and activation layers with pre-trained weights as parameters. It can perform classification prediction on a massive amount of raw material in the original material library, obtaining the classification probability distribution corresponding to each raw material, i.e., at least one predicted probability of at least one structure corresponding to each raw material.
[0142] In some embodiments, Figure 8 The convolutional layers of the CNN classification network can be MobileNet or other types of convolutional network layers. The specific choice depends on the actual situation, and this application does not limit the choice.
[0143] S002. When the largest prediction probability among at least one prediction probability is greater than a preset probability threshold, perform image content recognition on the original material to obtain the content tag corresponding to the original material.
[0144] In this embodiment, when the maximum predicted probability of the original material relative to at least one structure, obtained through classification prediction, is greater than a preset probability threshold, it indicates that the original material is similar in shape to the text structure and can be used as a preset material to replace the text structure. The image processing device can further perform image content recognition on the original material to obtain the content tag corresponding to the original material.
[0145] In some embodiments, the image processing apparatus may utilize artificial intelligence-based image recognition methods to identify the content contained in the original material. For example... Figure 9 As shown, from the salmon sushi image, we can identify content such as "Japanese food," "sushi," and "salmon." The image processing device can then label the original image with corresponding content tags based on the identified content. For example, by using an image recognition model to perform content recognition on the image, we can obtain the recognized content in the following code form:
[0146] {“Response”:{
[0147] “Labels”:{
[0148] {
[0149] "Name": "Tower", / / Material Name
[0150] "FirstCategory": "Scene", / / First-level directory
[0151] "SecondCategory": "Architecture", / / Second-level directory
[0152] "Confidence": 81 / / Confidence level, indicating the probability that the material is about a tower.
[0153] },
[0154] {
[0155] “Name”: “Night”,
[0156] "FirstCategory": "Scene"
[0157] "SecondCategory": "Natural Scenery"
[0158] "Confidence": 79
[0159] },
[0160] {
[0161] “Name”: “Skyline”,
[0162] "FirstCategory": "Scene"
[0163] "SecondCategory": "Natural Scenery"
[0164] "Confidence": 77
[0165] },
[0166] }
[0167] In some embodiments, before the image processing device assigns corresponding content tags to the original material based on the identified content, it may first deduplicate the identified content and mark the confidence level corresponding to each identified content in the content tag, such as ((Tower, 81), (Night, 79)...) etc.
[0168] S003. Take at least one predicted probability of at least one structure corresponding to the original material as structural information, and store the original material and its corresponding structural information and content tags into the preset material library.
[0169] It is understood that in this embodiment of the application, by automatically filtering materials suitable for replacing text structures and identifying the structural information and content tags of the materials, a material library is constructed. This can accumulate materials, greatly improve the reusability of materials, and facilitate rapid matching based on the structural information and content tags of each material, thereby improving the generation efficiency of special effects text images.
[0170] In some embodiments, based on Figure 7The image processing apparatus can further perform S201-S202, which will be described in combination with the respective steps.
[0171] S201, obtain at least one of material color information of the original material and an adaptive region of the original material.
[0172] In the embodiment of the application, the image processing apparatus can obtain at least one of RGB channel, grayscale or brightness information of the original material as the material color information; and / or, the image processing apparatus can obtain the adaptive region of the original material in the text according to the structure information of the original material, so as to obtain at least one of the material color information of the original material and the adaptive region of the original material in the text.
[0173] For example, for the structure "day", it can be located in the lower region in the "spring" character, or it can be located in the upper region in the "gnomon" character, so for the original material related to the structure "day", the image processing apparatus can identify the display region of the original material in the text according to the specific shape details and other information, so as to obtain the adaptive region of the original material.
[0174] S202, store at least one of the material color information and the adaptive region, the original material, the corresponding structure information and the content label into a preset material library.
[0175] In the embodiment of the application, the image processing apparatus can store at least one of the material color information and the adaptive region, the original material, the corresponding structure information and the content label into a preset material library, so that subsequent matching of the materials in the preset material library can be performed according to at least one of the color dimension, the position dimension, the shape dimension and the content dimension.
[0176] It can be understood that in the embodiment of the application, the information of the preset material library can be further enriched from the color dimension and the position dimension, so as to improve the flexibility of material matching in the preset material library.
[0177] In some embodiments, referring to Figure 10 , Figure 10 is an optional flow diagram of an image processing method provided by the embodiment of the application, based on Figure 5 S104 can be implemented by performing S301-S304, which will be described in combination with the respective steps.
[0178] S301, determine a set of to-be-matched materials matched with the iconic structure from at least one material.
[0179] In the embodiments of the present application, the structure information represents at least one prediction probability of each material corresponding to at least one structure in a preset structure category. The image processing device can determine, according to the iconic structure, the prediction probability corresponding to the iconic structure in the structure information of each material in the preset material library, and use at least one material with the prediction probability corresponding to the iconic structure being greater than the preset structure matching threshold as the set of materials to be matched.
[0180] Exemplarily, when the iconic structure is the stroke "丿" or "日", the image processing device can use, in the preset material library, at least one material with the prediction probability relative to the structure of "丿" or "日" being greater than 50% as the set of materials to be matched.
[0181] S302. For each material to be matched in the set of materials to be matched, calculate the structure score corresponding to each material to be matched according to the preset structure matching weight and the prediction probability in the structure information corresponding to each material to be matched.
[0182] In the embodiments of the present application, the prediction probability corresponding to the iconic structure in the structure information of each material to be matched may be different. The image processing device can calculate the structure score corresponding to each material to be matched according to the prediction probability corresponding to the iconic structure in the structure information of each material to be matched and the preset structure matching weight.
[0183] S303. Combine the preset semantic matching weight to calculate the similarity between the content label of each material to be matched and the extended semantic information, and obtain the semantic score corresponding to each material to be matched.
[0184] In the embodiments of the present application, the image processing device can calculate the similarity value, that is, the similarity, between the content label of each material to be matched and the extended semantic information according to the vector similarity calculation method. The image processing device combines the similarity with the preset semantic matching weight to obtain the structure score corresponding to each material to be matched.
[0185] In some embodiments, the content label of each material contains at least one recognition content, and each recognition content contains the confidence corresponding to the recognition content as the label confidence. For each material to be matched, the image processing device can determine, from at least one content label, a candidate content label with the label confidence being greater than or equal to the preset confidence threshold; calculate the similarity between the candidate content label and the extended semantic information, and obtain the semantic score corresponding to each material to be matched. That is, the image processing device can screen out the recognition content with high confidence from at least one recognition content for calculating the semantic score to improve the calculation efficiency. Exemplarily, the preset confidence threshold can be 50%, or can be preset to other values, and is specifically selected according to the actual situation, which is not limited in the embodiments of the present application.
[0186] In some embodiments, the preset semantic matching weight can be greater than the preset structure matching weight, or other settings can be made according to actual conditions, and the embodiments of the present application are not limited.
[0187] In S304, the special effect material is filtered from the set of to-be-matched materials based on the structure score and the semantic score.
[0188] In the embodiments of the present application, the image processing apparatus can perform comprehensive filtering based on the obtained structure score and semantic score to filter the special effect material from the set of to-be-matched materials.
[0189] In some embodiments, the image processing apparatus can perform summation or average calculation on the structure score and the semantic score to obtain a comprehensive score, and take the to-be-matched material with the highest comprehensive score in the set of to-be-matched materials as the special effect material.
[0190] In some embodiments, the image processing apparatus can also take at least one to-be-matched material with a comprehensive score greater than a preset total score threshold as a candidate material after obtaining the comprehensive score, and perform secondary filtering on the candidate material according to the specific values of the structure score and the semantic score, such as selecting the candidate material with the highest semantic score as the special effect material, and the like.
[0191] Here, it should be noted that when the number of to-be-matched materials, candidate content labels or candidate materials filtered by the image processing apparatus according to the above-mentioned preset structure matching threshold, preset confidence threshold and preset total score threshold is too small to meet the filtering requirements, the above-mentioned preset structure matching threshold, preset confidence threshold and preset total score threshold can be dynamically relaxed through a dynamic relaxation process similar to S102 to ensure more accurate filtering results.
[0192] In some embodiments, the image processing apparatus can also filter the special effect material from the set of to-be-matched materials based on the structure score and the semantic score, in combination with the material color information and the adaptation region of the material. The above-mentioned filtering method can be combined, selected or deformed according to actual conditions, and the embodiments of the present application are not limited.
[0193] It can be understood that in the embodiments of the present application, the use of landmark structures and extended semantic information for automatic matching in the preset material library covers a large amount of materials in the preset material library in material selection, and improves the efficiency and richness of the special effect text image generation.
[0194] In some embodiments, referring to Figure 11 , Figure 11 is an optional flowchart of the image processing method provided by the embodiments of the present application, based on Figure 5 , S104 can also be implemented by performing S401-S403, which will be described in combination with each step.
[0195] S401、According to the prediction probability of the corresponding landmark structure in the structure information of each material in the preset material library, and in combination with the preset structure matching weight, a structure score corresponding to each material is calculated.
[0196] S402、In combination with the preset semantic matching weight, the similarity of the content label and the extended semantic information of each material is calculated to obtain a semantic score corresponding to each material.
[0197] In the embodiments of the present application, for each material in the preset material library, the image processing device can calculate a structure score corresponding to each material according to the prediction probability of the landmark structure in the structure information of each material, in combination with the preset structure matching weight; at the same time, in combination with the preset semantic matching weight, the similarity of the content label and the extended semantic information of each material is calculated in the full set range of the preset material library to obtain a semantic score corresponding to each material.
[0198] In the embodiments of the present application, the process of calculating the structure score and the semantic score by the image processing device is similar to the process in S301 and S302, which will not be described here.
[0199] S403、Based on the structure score and the semantic score, a special effect material is selected from the preset material library.
[0200] In the embodiments of the present application, the image processing device can sum or average the structure score and the semantic score to obtain a comprehensive score corresponding to each material in the preset material library; and based on the comprehensive score, a special effect material is selected from the preset material library. S403 is to select a special effect material through a similar process in S304 in the full set range of the preset material library, which will not be described here.
[0201] In some embodiments, based on the above S201-S203, S403 can be implemented by performing S4031-S4033, which will be described in combination with each step.
[0202] S4031、According to the structure score and the semantic score, a candidate material set is selected from the preset material library.
[0203] In the embodiments of the present application, based on at least one of the material color information obtained in the above S401-S403 and the adaptation region, when the image processing device performs screening in the preset material library, the candidate material set corresponding to the preliminary screening can be obtained according to the structure score and the semantic score.
[0204] S4032、Obtain the color information corresponding to the landmark structure, or obtain the background color information of the to-be-processed text image as the to-be-matched color information.
[0205] In the embodiments of the present application, the image processing apparatus can obtain color information corresponding to the landmark structure, or obtain background color information of the to-be-processed text image as the to-be-matched color information, so that the special effect material matched through the to-be-matched color information can be consistent with the color of the corresponding text or the whole to-be-processed text image.
[0206] S4033、According to at least one of the adaptation region of each candidate material and the material color information, a candidate feature material that matches at least one of the regions occupied by the landmark structure is selected from the candidate material set as the special effect material.
[0207] In the embodiments of the present application, the image processing apparatus can further screen the candidate material set according to at least one of the adaptation region of each candidate material and the material color information, and select a candidate feature material that matches at least one of the regions occupied by the landmark structure as the special effect material.
[0208] It can be understood that the image processing apparatus can further screen the material from the dimensions of color and position to obtain the special effect material, so as to further improve the performance effect of generating the text special effect image according to the special effect material.
[0209] In the following, an exemplary application of the embodiments of the present application in an actual application scenario will be described. As shown in Figure 12 The exemplary application process of the embodiments of the present application in an actual application scenario can include a main process and an auxiliary process, wherein the main process includes:
[0210] S501, receiving a to-be-processed text image.
[0211] In S501, a text image that needs to be converted into a special effect word input by a user is received as a to-be-processed text image.
[0212] S502, training a key stroke annotator and identifying key strokes through the key stroke annotator.
[0213] In S502, the key stroke annotator, i.e., a stroke recognition model, is trained, and the pre-trained key stroke annotator is used to perform stroke disassembly and recognition on each text included in the to-be-processed text image to obtain at least one stroke, i.e., at least one text structure. Here, the execution process of S502 is consistent with that described in S101, and will not be described here again.
[0214] S503, filtering and selecting key strokes.
[0215] In S503, at least one stroke recognized is filtered according to a preset screening rule, and a key stroke, i.e., a landmark text structure, is selected. Here, the execution process of S503 is consistent with that described in S102, and will not be described here again.
[0216] S504, performing word segmentation on the to-be-processed character image.
[0217] In S504, word segmentation is performed on the characters in the to-be-processed image, to obtain at least one word.
[0218] S505, expanding word meaning.
[0219] In S505, word meaning expansion is performed on the at least one word obtained by word segmentation according to a word vector similarity calculation method, to obtain expanded semantic information. The execution processes of S504 and S505 are consistent with those described in S103, and thus will not be described here.
[0220] S506, matching materials.
[0221] In S506, materials are automatically matched in a material library according to the key strokes and the expanded semantic information, to obtain special effect materials. The execution process of S506 is consistent with that described in S104, and thus will not be described here.
[0222] S507, placing special effect materials.
[0223] In S507, the special effect materials are placed in the image region corresponding to the key strokes, to cover the key strokes, and a special effect character image is obtained.
[0224] S508, outputting the special effect character image.
[0225] The execution processes of S507 and S508 are consistent with those described in S105, and thus will not be described here.
[0226] The embodiments of the present application show a construction method of a preset material library in the auxiliary process of Figure 12 The construction method of the preset material library is as follows:
[0227] S601, obtaining a large amount of original materials.
[0228] In S601, a large amount of original materials can be obtained from copyright libraries and material pictures accumulated by designers.
[0229] S602, identifying material content.
[0230] In S602, an image recognition model of artificial intelligence is used to identify the material content of the original materials, and at least one content label is added to each original material according to the identification result. The execution process of S602 is consistent with that described in S002, and thus will not be described here.
[0231] S603, classifying material shapes.
[0232] In S603, the original material is classified according to the shape of the material by the pre-trained structure classification model, and at least one probability corresponding to at least one structure of each original material is obtained. Here, the execution process of S603 is consistent with that described in S001, and will not be repeated here.
[0233] In S604, a preset material library is constructed.
[0234] In S604, the original material with the content label and at least one probability corresponding to at least one structure is stored, thereby constructing a preset material library. Here, the execution process of S604 is consistent with that described in S003, and will not be repeated here.
[0235] It can be understood that in the embodiments of the present application, the key strokes and extended semantic information are automatically matched and selected in the material library to select special effect materials, which can greatly improve the reuse rate of materials, solve the problem that artificial energy is limited and cannot be involved in massive materials in design, and through key stroke recognition and special effect material covering of the to-be-processed text image, the special effect text image is automatically generated by the machine, which greatly reduces the design cost and improves the generation efficiency of the special effect text image. And by using the extended multiple word meanings and the massive material library, the strokes are automatically selected and the materials are automatically matched and replaced by the machine, so that more rich and diverse special effect materials can be selected to generate special effect text images on the basis of meeting the design specifications, and the performance effect of the special effect text image is improved. And by constructing the material library, the materials can be deposited and the reuse degree of the materials can be greatly improved.
[0236] The following continues to describe an example structure of the image processing apparatus 255 provided by the embodiments of the present application as a software module. In some embodiments, as shown in Figure 4 The software module stored in the image processing apparatus 255 of the memory 250 can include:
[0237] The structure recognition module 2551 is configured to perform text structure disassembly on the obtained to-be-processed text image based on a preset structure category, and obtain at least one text structure.
[0238] The screening module 2552 is configured to screen the at least one text structure according to a preset screening rule, and obtain a landmark structure.
[0239] The semantic module 2553 is configured to obtain semantic information in the to-be-processed text image and perform semantic extension to obtain extended semantic information.
[0240] The matching module 2554 is configured to match the landmark structure and the extended semantic information with structure information and a content label corresponding to each material in at least one material in a preset material library, and obtain a special effect material.
[0241] cover the landmark structure in the to-be-processed text image using the special effect material, to obtain a special effect text image.
[0242] In some embodiments, each of the preset screening rules corresponds to a preset rule priority, and the preset screening rules include at least one of the following:
[0243] The landmark structure occupies an area proportion of the to-be-processed text image, and the area proportion is greater than or equal to a preset area proportion threshold;
[0244] The number of landmark structures is less than or equal to a first preset number threshold;
[0245] The landmark structure is a preset structure in the preset structure category;
[0246] The landmark structure is located in a preset region in the to-be-processed text image.
[0247] In some embodiments, the image processing apparatus further includes an adjustment module, which is configured to adjust the preset screening rules to obtain an adjusted screening rule when the number of landmark structures is less than a second preset number threshold; the second preset number threshold is less than the first preset number threshold; the adjusted screening rule is used for dynamic relaxation processing of the preset screening rules to increase the number of landmark structures obtained according to the adjusted screening rule; and the at least one text structure is re-screened according to the adjusted screening rule to obtain the landmark structure; wherein the dynamic relaxation processing includes at least one of reducing the number of used rules and reducing the preset area proportion threshold in the preset screening rule.
[0248] In some embodiments, the semantic module 2553 is further configured to perform text content recognition on the to-be-processed text image to obtain a character sequence corresponding to the to-be-processed text image; perform word segmentation processing on the character sequence to obtain at least one word; perform word meaning extension on the at least one word to obtain at least one extended word corresponding to each word; and take each word and its corresponding at least one extended word as the extended semantic information.
[0249] In some embodiments, the semantic module 2553 is further configured to, for each word in the at least one word, calculate a similarity between the each word and each preset word vector in a preset word vector library; and take a preset word vector with a similarity greater than or equal to a preset similarity threshold as an extended word corresponding to the each word, to obtain the at least one extended word.
[0250] In some embodiments, the matching module 2554 is further configured to determine, from the at least one material, a set of to-be-matched materials matching the landmark structure; for each to-be-matched material in the set of to-be-matched materials, calculate a structure score of the each to-be-matched material according to a preset structure matching weight and a prediction probability in structure information corresponding to the each to-be-matched material; calculate a similarity between a content label of the each to-be-matched material and the extended semantic information to obtain a semantic score of the each to-be-matched material in combination with a preset semantic matching weight; and filter the special effect material from the set of to-be-matched materials based on the structure score and the semantic score.
[0251] In some embodiments, the content label is at least one recognition content, and each recognition content includes a label confidence; the matching module 2554 is further configured to, for the each to-be-matched material, determine a candidate content label with a label confidence greater than or equal to a preset confidence threshold from the at least one content label; and calculate a similarity between the candidate content label and the extended semantic information to obtain the semantic score of the each to-be-matched material.
[0252] In some embodiments, the matching module 2554 is further configured to calculate a structure score of each material in the preset material library according to a prediction probability in structure information of the each material corresponding to the landmark structure in combination with a preset structure matching weight; calculate a similarity value between a content label of the each material and the extended semantic information to obtain a semantic score of the each material in combination with a preset semantic matching weight; and filter the special effect material from the preset material library based on the structure score and the semantic score.
[0253] In some embodiments, the image processing apparatus further comprises a material processing module, which is configured to, before the matching in the preset material library based on the landmark structure and the extended semantic information to obtain the special effect material, perform classification prediction on original materials according to the preset structure category to obtain the at least one prediction probability of the at least one structure corresponding to the original materials; when a maximum prediction probability in the at least one prediction probability is greater than a preset probability threshold, perform image content recognition on the original materials to obtain a content label corresponding to the original materials; take the at least one prediction probability of the at least one structure corresponding to the original materials as the structure information; and store the material color information and at least one of the adaptation regions, the original materials, the structure information corresponding to the original materials, and the content label in the preset material library.
[0254] In some embodiments, the material processing module is further configured to acquire material color information of the original material and at least one of the fitting area of the original material; and store the original material and the corresponding material color information, the at least one of the fitting area, the structure information and the content label in the preset material library.
[0255] In some embodiments, the matching module 2554 is further configured to filter a candidate material set from the preset material library according to the structure score and the semantic score; acquire color information corresponding to the landmark structure or background color information of the to-be-processed text image as to-be-matched color information; and filter, from the candidate material set, a candidate material matching at least one of the areas occupied by the landmark structure as the special effect material according to at least one of the fitting area and the material color information of each candidate material.
[0256] In some embodiments, the matching module 2554 is further configured to perform summation or average calculation on the structure score and the semantic score to obtain a comprehensive score corresponding to each material in the preset material library; and filter the special effect material from the preset material library based on the comprehensive score.
[0257] In some embodiments, the material processing module is further configured to, for each structure in the preset structure category, acquire a preset sample number of original materials corresponding to the each structure from an original material library to obtain a sample material set; perform model training on an initial multi-classification neural network by using the sample material set to obtain a structure classification model; and perform classification on remaining original materials in the original material library by using the structure classification model to obtain the at least one predicted probability of the at least one structure corresponding to the remaining original materials, thereby completing the classification of the original materials.
[0258] In some embodiments, the structure recognition module 2551 is further configured to perform target detection on the to-be-processed text image based on the preset structure category to predict a plurality of image areas corresponding to the preset structure category; each image area contains at least one region confidence; the at least one region confidence represents at least one probability that the each image area corresponds to at least one structure in the preset structure category; and a structure corresponding to a region confidence greater than or equal to a preset structure confidence threshold is taken as the at least one text structure in the to-be-processed text image.
[0259] In some embodiments, each image region contains position information; the overlay module 2555 is further configured to perform image preprocessing on the special effect material to obtain preprocessed special effect material; and overlay the preprocessed special effect material on the image region occupied by the landmark structure according to the position information to obtain the special effect text image; wherein the image preprocessing includes at least one of scaling processing and rotation processing; and the scaling processing is used to adjust the size of the special effect material according to the image region occupied by the landmark structure.
[0260] In some embodiments, the matching module 2554 is further configured to match the landmark structure and the extended semantic information with the structure information and the content label corresponding to each of at least one material in a preset material library to obtain a plurality of alternative materials; and use any of the plurality of alternative materials as the special effect material.
[0261] The image processing apparatus further includes a replacement module, which is configured to, after the landmark structure in the text image to be processed is overlaid with the special effect material to obtain a special effect text image, generate a new special effect text image according to the remaining materials in the plurality of alternative materials when a special effect text image replacement instruction is received.
[0262] It should be noted that the above description of the device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects to the method embodiments. For technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.
[0263] The embodiments of the present application provide a computer readable storage medium storing executable instructions, wherein the executable instructions, when executed by a processor, will cause the processor to perform the method provided by the embodiments of the present application, for example, the method shown in Figure 5 、 7 , 11, 12.
[0264] In some embodiments, the computer readable storage medium can be FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM, etc. memory; or can be various devices including one or any combination of the above memories.
[0265] In some embodiments, the executable instructions can be in the form of programs, software, software modules, scripts or codes, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and can be deployed in any form, including being deployed as independent programs or as modules, components, subroutines or other units suitable for use in a computing environment.
[0266] By way of example, executable instructions can correspond to a file in a file system, can be stored in a portion of a file that holds other programs or data, for example, one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, for example, files that store one or more modules, subprograms, or portions of code.
[0267] By way of example, executable instructions can be deployed to be executed on one computer, or on multiple computers of a system, or on multiple computers distributed among multiple locations and interconnected by a communication network.
[0268] To sum up, the embodiment of the present application realizes the automatic recognition of the representative structure such as strokes in the text by carrying out the text structure disassembly in the text image to be processed and automatically screening the symbolic structure from the at least one text structure obtained by the disassembly, and further matches the special effect material from the preset material library based on the recognized symbolic structure and in combination with the extended semantic information, thereby greatly saving the labor cost, greatly improving the generation efficiency of the special effect text image, and being capable of screening more diverse special effect materials by the extension of the semantic information, and improving the performance effect of the special effect text image. Moreover, by using the extended multiple word meanings and the massive material library, the strokes are selected and the material is matched and replaced by the machine, the more diverse special effect materials can be selected to generate the special effect text image on the basis of guaranteeing the compliance with the design specification, the performance effect of the special effect text image is improved. Furthermore, by constructing the material library, the materials can be deposited and the reuse degree of the materials is greatly improved.
[0269] The above merely describes the embodiments of the present application, but is not used to limit the protection scope of the present application. Any modification, equivalent replacement, and improvement within the spirit and scope of the present application shall be included in the protection scope of the present application.
Claims
1. An image processing method, characterized by, The method comprises: based on a preset structure category, the obtained text image is text structure disassembled, and at least one text structure is obtained; According to the preset screening rule, the at least one text structure is screened to obtain a landmark structure; Obtain semantic information in the text image to be processed and perform semantic expansion to obtain expanded semantic information; Based on the landmark structure and the expanded semantic information, the structure information and the content label corresponding to each material in at least one material in the preset material library are matched to obtain a special effect material; Using the special effect material to cover the landmark structure in the text image to be processed, a special effect text image is obtained.
2. The method of claim 1, wherein, Each rule in the preset screening rule corresponds to a preset rule priority, and the preset screening rule comprises at least one of the following: The area proportion of the landmark structure in the text image to be processed is greater than or equal to a preset area proportion threshold; The number of landmark structures is less than or equal to a first preset number threshold; The landmark structure is a preset structure in the preset structure category; The landmark structure is located in a preset area in the text image to be processed.
3. The method of claim 2, wherein, The method further comprises: When the number of landmark structures is less than a second preset number threshold, the preset screening rule is adjusted to obtain an adjusted screening rule; the second preset number threshold is less than the first preset number threshold; the adjusted screening rule is used for dynamic relaxation processing of the preset screening rule to increase the number of landmark structures obtained according to the adjusted screening rule; According to the adjusted screening rule, the at least one text structure is re-screened to obtain the landmark structure; Wherein, the dynamic relaxation processing includes at least one of reducing the number of rules used and reducing the preset area proportion threshold in the preset screening rule.
4. The method of claim 1, wherein, The method for obtaining semantic information in the text image to be processed and performing semantic expansion to obtain expanded semantic information comprises: Text content recognition is performed on the text image to be processed to obtain a character sequence corresponding to the text image to be processed; The character sequence is segmented to obtain at least one word; The at least one word is extended to obtain at least one extended word corresponding to each word; Each word and its corresponding at least one extended word are used as the expanded semantic information; Wherein, the at least one word is extended to obtain at least one extended word, comprising: For each word in the at least one word, the similarity between each word and each preset word vector in a preset word vector library is calculated; The preset word vector with a similarity greater than or equal to a preset similarity threshold is used as the extended word corresponding to each word to obtain the at least one extended word.
5. The method of claim 1, wherein, The method for matching the structure information and the content label corresponding to each material in at least one material in the preset material library based on the landmark structure and the expanded semantic information to obtain a special effect material comprises any one of the following: From the at least one material, a set of to-be-matched materials matching the landmark structure is determined; For each to-be-matched material in the to-be-matched material set, a structure score corresponding to each to-be-matched material is calculated according to a preset structure matching weight and a prediction probability in structure information corresponding to each to-be-matched material; In combination with a preset semantic matching weight, a similarity between a content label of each to-be-matched material and the extended semantic information is calculated to obtain a semantic score corresponding to each to-be-matched material; The special effect material is filtered from the to-be-matched material set based on the structure score and the semantic score; Or, According to a prediction probability in structure information corresponding to each material in the preset material library and in combination with a preset structure matching weight, a structure score corresponding to each material is calculated; In combination with a preset semantic matching weight, a similarity between a content label of each material and the extended semantic information is calculated to obtain a semantic score corresponding to each material; The special effect material is filtered from the preset material library based on the structure score and the semantic score.
6. The method of claim 5, wherein, The content label is at least one recognition content, and each recognition content includes a label confidence; The combination of the preset semantic matching weight, the calculation of the similarity between the content label of each to-be-matched material and the extended semantic information, and the obtaining of the semantic score corresponding to each to-be-matched material includes: For each to-be-matched material, a candidate content label with a label confidence greater than or equal to a preset confidence threshold is determined from the at least one content label; The similarity between the candidate content label and the extended semantic information is calculated to obtain the semantic score corresponding to each to-be-matched material.
7. The method according to any one of claims 5-6, characterized in that, Before the matching of the structure information and the content label of each material in at least one material in the preset material library based on the landmark structure and the extended semantic information, the method further includes: According to the preset structure category, the original material is classified and predicted to obtain the at least one prediction probability of the at least one structure corresponding to the original material; When the maximum prediction probability in the at least one prediction probability is greater than a preset probability threshold, image content recognition is performed on the original material to obtain the content label corresponding to the original material; The at least one prediction probability of the at least one structure corresponding to the original material is taken as the structure information, and the original material, the structure information corresponding thereto, and the content label are stored in the preset material library.
8. The method of claim 7, wherein, The method further includes: Obtaining material color information of the original material and at least one of an adaptation region of the original material; The material color information and the at least one of the adaptation region are stored in the preset material library together with the original material, the structure information corresponding thereto, and the content label; The filtering of the special effect material from the preset material library based on the structure score and the semantic score includes: According to the structure score and the semantic score, a candidate material set is filtered from the preset material library. Obtaining color information corresponding to the landmark structure, or obtaining background color information of the to-be-processed text image as to-be-matched color information; According to at least one of the fitting area and the material color information of each candidate material, the candidate material set is filtered to obtain a candidate material that matches at least one of the areas occupied by the landmark structure as the special effect material.
9. The method of claim 7, wherein, The method comprises the following steps of: For each structure in the preset structure category, a preset sample number of original materials corresponding to each structure is obtained from an original material library to obtain a sample material set; Using the sample material set, an initial multi-classification neural network is trained to obtain a structure classification model; Using the structure classification model, the remaining original materials in the original material library are classified to obtain at least one prediction probability of the at least one structure corresponding to the remaining original materials, thereby completing the classification of the original materials.
10. The method of claim 1, wherein, The method comprises the following steps of: Based on the preset structure category, target detection is performed on the to-be-processed text image to predict a plurality of image regions corresponding to the preset structure category; each image region contains at least one region confidence; the at least one region confidence represents at least one probability that each image region corresponds to at least one structure in the preset structure category; The structure corresponding to the region confidence greater than or equal to a preset structure confidence threshold is taken as the at least one text structure in the to-be-processed text image.
11. The method of claim 10, wherein, Each image region contains position information; the method comprises the following steps of: Image preprocessing is performed on the special effect material to obtain preprocessed special effect material; According to the position information, the preprocessed special effect material is overlaid on the image region occupied by the landmark structure to obtain the special effect text image; wherein the image preprocessing comprises: At least one of scaling processing and rotation processing; the scaling processing is used to adjust the size of the special effect material according to the image region occupied by the landmark structure.
12. The method according to any one of claims 1 to 6, characterized in that, The method comprises the following steps of: The structure information and content label corresponding to each material in at least one material in the preset material library are matched with the landmark structure and the extended semantic information to obtain a special effect material, comprising: The structure information and content label corresponding to each material in at least one material in the preset material library are matched with the landmark structure and the extended semantic information to obtain a plurality of alternative materials; Any alternative material in the plurality of alternative materials is taken as the special effect material; After the landmark structure is overlaid with the special effect material to obtain the special effect text image, the method further comprises: When receiving the special effect text image replacement instruction, a new special effect text image is generated according to the remaining materials in the plurality of alternative materials.
13. An image processing apparatus characterized by comprising: Comprise: The structure identification module is used for carrying out text structure disintegration on the obtained to-be-processed text image based on a preset structure category, and obtaining at least one text structure. The screening module is used for screening the at least one text structure according to a preset screening rule, and obtaining a landmark structure. The semantic module is used for obtaining semantic information in the to-be-processed text image and performing semantic expansion, and obtaining expanded semantic information. The matching module is used for matching the landmark structure and the expanded semantic information with structure information and content labels corresponding to each material in at least one material in a preset material library, and obtaining a special effect material. The covering module is used for covering the landmark structure in the to-be-processed text image with the special effect material, and obtaining a special effect text image.
14. The apparatus of claim 13, wherein, Each rule in the preset screening rule corresponds to a preset rule priority, and the preset screening rule comprises at least one of the following: The area proportion of the landmark structure in the to-be-processed text image is greater than or equal to a preset area proportion threshold value; and the number of landmark structures is less than or equal to a first preset number threshold value; The landmark structure is a preset structure in the preset structure category; The landmark structure is located in a preset area in the to-be-processed text image.
15. The apparatus of claim 14, wherein, Further comprising an adjusting module for: When the number of landmark structures is less than a second preset number threshold value, adjusting the preset screening rule to obtain an adjusted screening rule; the second preset number threshold value is less than the first preset number threshold value; and the adjusted screening rule is used for dynamically relaxing the preset screening rule to increase the number of landmark structures obtained according to the adjusted screening rule; According to the adjusted screening rule, the at least one text structure is re-screened to obtain the landmark structure; wherein the dynamic relaxation processing comprises at least one of reducing the number of used rules and reducing the preset area proportion threshold value in the preset screening rule.
16. The apparatus of claim 13, wherein, The semantic module is further used for: Carrying out text content recognition on the to-be-processed text image to obtain a character sequence corresponding to the to-be-processed text image; carrying out word segmentation processing on the character sequence to obtain at least one word; carrying out word meaning expansion on the at least one word to obtain at least one expanded word corresponding to each word; and taking each word and its corresponding at least one expanded word as the expanded semantic information; wherein the word meaning expansion on the at least one word segmentation to obtain at least one expanded word comprises: for each word in the at least one word, calculating the similarity between the each word and each preset word vector in a preset word vector library; and taking a preset word vector with a similarity greater than or equal to a preset similarity threshold value as an expanded word corresponding to the each word to obtain the at least one expanded word.
17. The apparatus of claim 13, wherein, The matching module is further used for: determining, from the at least one material, a set of to-be-matched materials matching the landmark structure; for each to-be-matched material in the set of to-be-matched materials, calculating a structure score corresponding to the each to-be-matched material according to a preset structure matching weight and a prediction probability in structure information corresponding to the each to-be-matched material; calculating similarity between a content label of the each to-be-matched material and the extended semantic information to obtain a semantic score corresponding to the each to-be-matched material in combination with a preset semantic matching weight; and screening the special effect material from the set of to-be-matched materials based on the structure score and the semantic score. Alternatively, calculating a structure score corresponding to the each material according to a prediction probability in structure information corresponding to the each material in the preset material library and in combination with a preset structure matching weight; calculating similarity between a content label of the each material and the extended semantic information to obtain a semantic score corresponding to the each material in combination with a preset semantic matching weight; and screening the special effect material from the preset material library based on the structure score and the semantic score. The content label is at least one recognition content, and each recognition content contains a label confidence; 18. The apparatus of claim 17, wherein, The matching module is further configured to: for the each to-be-matched material, determining a candidate content label with a label confidence greater than or equal to a preset confidence threshold from the at least one content label; and calculating similarity between the candidate content label and the extended semantic information to obtain the semantic score corresponding to the each to-be-matched material. comprising:
19. An electronic device, comprising: a memory configured to store executable instructions; a processor configured to execute the executable instructions stored in the memory to implement the method in any one of claims 1 to 12. executable instructions stored in the memory, and configured to be executed by the processor to implement the method in any one of claims 1 to 12.
20. A computer-readable storage medium, characterized in that,
Citation Information
Patent Citations
Image processing method and system
CN105279186A
Method and system for converting Chinese character fonts in image, computer equipment and medium
CN110135530A