Image Processing Method, Apparatus, Device, Readable Storage Medium and Program Product

By extracting the to-processed area from the original image and determining the control point for deformation processing, the target video with dynamic switching is generated, and the problem of high cost and low efficiency of animation video production is solved, and efficient and low-cost animation video generation is achieved.

CN115115754BActive Publication Date: 2025-07-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210684717.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-16
Publication Date
2025-07-25
Estimated Expiration
2042-06-16

AI Technical Summary

Technical Problem

The prior art has high cost and low efficiency when producing animation videos, making it difficult to adapt to a wide variety of business scenarios, and there are limited types of animation videos produced manually.

Method used

By extracting the image of the area to be processed from the original image, determining the source control point and the target control point, deforming the target object, generating the target video, and dynamic switching of the target object between different forms.

Benefits of technology

It improves the production efficiency of animation videos, reduces production costs, and increases the types of animation videos, improves the attractiveness of target objects in the videos, and brings better advertising conversion effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115115754B_ABST
    Figure CN115115754B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides an image processing method, apparatus, device, readable storage medium, and program product, which can be applied to fields or scenarios such as artificial intelligence, vehicle-mounted scenarios, intelligent transportation, and image processing. The method includes: extracting an image of a region to be processed from an original image, where the image of the region to be processed includes a target object, and the target object is displayed in a first form in the image of the region to be processed; determining source control points of the image of the region to be processed, and determining target control points corresponding to the source control points; performing a deformation process on the image of the region to be processed according to the source control points and the target control points to obtain a deformed region image; the target object is displayed in a second form in the deformed region image; generating a target video according to the original image and the deformed region image, and the target object in the target video dynamically switches between the first form and the second form. Through the embodiment of the present application, an animated effect video can be automatically generated, the production efficiency of the animated effect video is improved, and the production cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technologies, and particularly to an image processing method, an image processing apparatus, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] The information display ability of static pictures is often limited. By dynamically displaying key parts in static pictures, the content to be conveyed in static pictures can be enhanced. For example, in advertising pictures, high-quality copywriting is the most important part of advertising materials. Through text motion effect videos, key information such as promotions, advertisements, and discounts can be expressed, attracting the attention of the target, thereby bringing better advertising conversion effects. Currently, when making text motion effect videos, for advertising materials in the initial production stage, designers need to make font motion effect templates and then perform post-production video production; for completed advertising materials, designers need to extract the text and then use tools to make text motion effect videos frame by frame. The above methods result in high production costs for motion effect videos and are not suitable for a wide variety of business scenarios. Moreover, making motion effect videos manually leads to few types of motion effect videos and low production efficiency of motion effect videos. Summary of the Invention

[0003] The present application provides an image processing method, apparatus, device, readable storage medium, and program product, which can automatically generate motion effect videos, improve the production efficiency of motion effect videos, and reduce production costs.

[0004] In a first aspect, the present application provides an image processing method, which includes:

[0005] Extract an image of a region to be processed from an original image, where the image of the region to be processed includes a target object, and the target object is displayed as a first form in the image of the region to be processed;

[0006] Determine source control points of the image of the region to be processed, and determine target control points corresponding to the source control points;

[0007] Perform a deformation process on the image of the region to be processed according to the source control points and the target control points to obtain a deformed region image; the target object is displayed as a second form in the deformed region image;

[0008] Generate a target video according to the original image and the deformed region image, and the target object in the target video dynamically switches between the first form and the second form.

[0009] In a second aspect, the present application provides an image processing apparatus, which includes:

[0010] An acquisition module, configured to extract an image of a region to be processed from an original image, where the image of the region to be processed includes a target object, and the target object is displayed as a first form in the image of the region to be processed;

[0011] A processing module, configured to determine source control points of the image of the region to be processed, and determine target control points corresponding to the source control points;

[0012] The processing module is further configured to perform a deformation process on the image of the region to be processed according to the source control points and the target control points to obtain a deformed region image; the target object is displayed as a second form in the deformed region image;

[0013] A generation module, configured to generate a target video according to the original image and the deformed region image, and the target object in the target video dynamically switches between the first form and the second form.

[0014] In a third aspect, the present application provides a computer device, including: a processor, a storage device, and a communication interface, the processor, the communication interface, and the storage device are connected to each other, wherein the storage device stores executable program codes, and the processor is configured to call the executable program codes to implement the above-mentioned image processing method.

[0015] In a fourth aspect, the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, the computer program includes program instructions, and the program instructions are executed by a processor to implement the above-mentioned image processing method.

[0016] In a fifth aspect, the present application provides a computer program product, where the computer program product includes a computer program or computer instructions, and the computer program or computer instructions are executed by a processor to implement the above-mentioned image processing method.

[0017] This application extracts the image of the area to be processed including the target object from the original image for deforming the target object, avoiding processing the entire image, reducing the computational amount, and thus improving the efficiency of the deformation process. By determining the source control points of the image of the area to be processed and their corresponding target control points, the terminal device can plan the approximate movement trajectory of the target object in the image of the area to be processed. The terminal device can perform the deformation process corresponding to the source control points to the target control points on the image of the area to be processed to obtain a deformed area image, and then generate a target video based on the original image and the deformed area image, thereby realizing the dynamic display of the target object between the first form and the second form, which can effectively enhance the attractiveness of the target object in the original image, attract the attention of the viewing object, and thus bring a better advertising conversion effect. Compared with manually making an animated video, this application can automatically generate an animated video based on the original graphics, improve the production efficiency of the animated video, and reduce the production cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 is a schematic diagram of the architecture of a data processing system provided by an exemplary embodiment of this application;

[0020] Figure 2 is a schematic flowchart of an image processing method provided by an exemplary embodiment of this application;

[0021] Figure 3A is a schematic diagram of the effect of an image processing method provided by an exemplary embodiment of this application;

[0022] Figure 3B is a schematic diagram of the effect of another image processing method provided by an exemplary embodiment of this application;

[0023] Figure 3C is a flowchart of an image processing provided by an exemplary embodiment of this application;

[0024] Figure 3D is a schematic flowchart of keyword image processing provided by an exemplary embodiment of this application;

[0025] Figure 3E is a schematic diagram of determining source control points provided by an exemplary embodiment of this application;

[0026] Figure 4It is a schematic flowchart of another image processing method provided by an exemplary embodiment of the present application;

[0027] Figure 5 It is a schematic diagram of the deformation of an image of an area to be processed provided by an exemplary embodiment of the present application;

[0028] Figure 6 It is a schematic block diagram of an image processing apparatus provided by an exemplary embodiment of the present application;

[0029] Figure 7 It is a schematic block diagram of a computer device provided by an exemplary embodiment of the present application. Specific embodiments

[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0031] It should be noted that the descriptions such as "first" and "second" involved in the embodiments of the present application are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Therefore, the technical features defined with "first" and "second" may explicitly or implicitly include at least one such feature.

[0032] The embodiments of the present invention can be applied to various fields or scenarios such as artificial intelligence, vehicle-mounted scenarios, intelligent transportation, and image processing. Next, several typical application fields or scenarios will be introduced.

[0033] An intelligent transportation system (ITS), also known as an intelligent transportation system (Intelligent Transportation System), effectively integrates advanced scientific and technological means (information technology, computer technology, data communication technology, sensor technology, electronic control technology, automatic control theory, operations research, artificial intelligence, etc.) into transportation, service control, and vehicle manufacturing, strengthening the connection among vehicles, roads, and users, thereby forming a comprehensive transportation system that ensures safety, improves efficiency, improves the environment, and saves energy.

[0034] Intelligent Vehicle Infrastructure Cooperative Systems (IVICS), referred to as IVICS, is a development direction of Intelligent Transportation Systems (ITS). IVICS uses advanced wireless communications and new generation Internet technologies to implement all-round dynamic real-time information interaction between vehicles and roads, and conducts active vehicle safety control and road cooperative management based on the collection and integration of dynamic traffic information in all time and space, fully realizing the effective coordination of people, vehicles and roads, ensuring traffic safety, and improving traffic efficiency, thus forming a safe, efficient and environmentally friendly road traffic system.

[0035] The present application can be applied to the above-mentioned fields. For example, by taking illegal vehicles in traffic images as target objects and dynamically displaying illegal vehicles, the efficiency of law enforcement personnel in detecting illegal behaviors can be improved. For another example, by dynamically displaying the traffic conditions ahead, speed limit information, etc. on the driver's car system, the violation rate can be reduced, thereby ensuring traffic safety.

[0036] Artificial Intelligence (AI) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that the machines have the functions of perception, reasoning and decision-making. Artificial intelligence technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level technology and software-level technology. Basic artificial intelligence technologies generally include technologies such as sensors, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technology mainly includes several major directions such as computer vision technology, machine learning / deep learning, etc. The solution provided in the embodiment of the present application involves machine learning, computer vision technology and other technologies under artificial intelligence technology, which will be described below.

[0037] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. Machine learning specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning. This application mainly involves the inductive learning technology in machine learning. The inductive learning technology aims to inductively extract general decision rules and patterns from a large amount of empirical data, and is a learning method that derives general rules from special cases. Specifically, the method proposed in this application removes the area where the target object is located in the original image, then uses an image restoration model to restore the original image after removal, and finally generates an animated video based on the text animation effect and the restored original image.

[0038] Computer Vision Technology (CV) Computer vision is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition and measurement, and further perform graphic processing to make the computer process images more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, intelligent transportation, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition. Specifically, the method proposed in this application can use, for example, a segmentation model to obtain the foreground area in the original image, perform text recognition processing and keyword matching processing on the text in the foreground area, finally obtain the keyword area, and then perform animation production on the keyword area.

[0039] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, vehicle networking, autonomous driving, 3D games, etc. With the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0040] Currently, when creating text effects, for the advertising materials in the initial production stage, designers are required to create font effect templates and then carry out post-production video production; for the completed advertising materials, designers need to extract the text and then use tools to create text effects frame by frame.

[0041] For the materials that have been produced, the above methods require secondary processing, resulting in high production costs, being difficult to scale up, and not being suitable for the huge number of advertising business scenarios. Moreover, by creating text effects manually, the variety of text effects depends on the number of templates created by designers, resulting in limited text effects and low production efficiency of text special effects.

[0042] This application provides a method for generating text dynamic effects. By recognizing the text in the advertising material picture, erasing the text and completing the background, and then making the text perform dynamic jumps such as simulating human body movements. Specifically, for the advertising materials in the initial production stage, this application can create text effects based on the tool for generating videos from pictures, which can improve production efficiency and increase the quantity scale of text effects; for the completed advertising materials, it can also be mass-produced. The method provided by this application can be used to erase the text in the picture and then reproduce it using the special effect of text jumping.

[0043] This application can simulate the movement law of the human body, making the jumping effect of the text more realistic and coordinated, effectively improving the attractiveness of the key core text information in the video generated from pictures, presenting a special effect of text jumping in the form of text expression, attracting the attention of the target, and thus bringing better advertising conversion effects.

[0044] It can be understood that in the specific implementation of this application, relevant data such as the original image and the target object are involved. When the above embodiments of this application are applied to specific products or technologies, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0045] This application will be specifically described through the following embodiments.

[0046] Please refer to Figure 1 , which is a schematic diagram of the architecture of a data processing system provided by an exemplary embodiment of this application. As shown in the figure, the data processing system may specifically include a terminal device 101 and a server 102. Among them, the terminal device 101 and the server 102 are connected through a network, for example, through a local area network, a wide area network, a mobile Internet, etc. The operation object operates on the browser or client application of the terminal device 101 to process various image data. The server 102 can respond to this operation and provide various services for image processing to the operation object.

[0047] Specifically, the terminal device 101 can obtain the image of the area to be processed in the original image, as well as the source control points and target control points of the image of the area to be processed, and then send the above data to the server 102; after obtaining the data, the server 102 performs deformation processing on the image of the area to be processed according to the source control points and target control points to obtain a deformed area image; the server 102 generates a target video for dynamically displaying the target object based on the original image and the deformed area image, and then sends the target video to the terminal device 101; finally, the target video is displayed on the terminal device 101.

[0048] The terminal device 101 is also referred to as a terminal, user equipment (UE), access terminal, user unit, mobile device, user terminal, wireless communication device, user agent, or user device. The terminal device can be a smart home appliance, a handheld device with wireless communication capabilities (such as a smart phone, a tablet computer), a computing device (such as a personal computer (PC), an in-vehicle terminal, a smart voice interaction device, a wearable device, or other intelligent devices), but is not limited thereto.

[0049] The server 102 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0050] It should be noted that the terminal device 101 and the server 101 described in this application can be the same computer device to implement the relevant methods in this application. For example, in the architecture of this application, only the terminal device 101 with computing capabilities is included. The terminal device 101 can obtain the image of the area to be processed in the original image, as well as the source control points and target control points of the image of the area to be processed, and then perform deformation processing on the image of the area to be processed according to the source control points and target control points to obtain a deformed area image, and generate a target video for dynamically displaying the target object based on the original image and the deformed area image, and finally display the target video on the terminal device 101.

[0051] It should be understood that the schematic diagram of the system architecture described in the embodiments of the present application is for more clearly illustrating the technical solutions of the embodiments of the present application, and does not constitute a limitation on the technical solutions provided by the embodiments of the present application. For example, the access method provided by the embodiments of the present application can be executed not only by server 102, but also by other servers or server clusters different from server 102 and capable of communicating with terminal device 101 and / or server 102. As is known to those of ordinary skill in the art, Figure 1 the number of terminal devices and servers in

[0052] is merely illustrative. According to the needs of business implementation, any number of terminal devices and servers can be configured. Moreover, with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems. In the subsequent embodiments, terminal device will be used to refer to the above-mentioned terminal device 101, and server will be used to refer to the above-mentioned server 102, and will not be elaborated in the subsequent embodiments. Figure 2 , Figure 2 Please refer to Figure 1 which is a schematic flowchart of an image processing method provided by an exemplary embodiment of the present application. Taking the application of this method to a terminal device (referring to the terminal device 101 in

[0053] S201. Extract an image of a region to be processed from the original image. The image of the region to be processed includes a target object, and the target object is displayed in a first form in the image of the region to be processed.

[0054] In the embodiments of the present application, the original image can be an image including a target object, and the image of the region to be processed is the image region where the target object is located. The original image can be, for example, an advertisement image, a product introduction image, etc., and the target object can be an object such as text, a commodity, etc. in the original image that the operator hopes to view. The target object is displayed in a first form in the image of the region to be processed, that is, the form before the target object is deformed. By extracting the target object from the original image, the terminal device facilitates the deformation processing of the target object, avoids processing the entire image, reduces the amount of calculation, and thus improves the efficiency of the deformation processing.

[0055] In one embodiment, since multiple characters generally appear as words or sentences, the target object can be the multiple characters that make up the words or sentences in the original image. The terminal device can use the multiple characters that make up the words or sentences as the target object. Then, the image region to be processed is the image region corresponding to the multiple characters that make up the words or sentences in the original image. By performing a deformation process on the image region corresponding to the multiple characters, it is avoided to process each character separately (at this time, each character will jump separately), which improves the overall coordination of the target object before and after deformation and enhances the viewing experience of the object.

[0056] S202. Determine the source control points of the image region to be processed, and determine the target control points corresponding to the source control points.

[0057] In the embodiments of the present application, the source control points are key control points selected in the image region to be processed for controlling the deformation process of the entire image region to be processed; the target control points are the position points corresponding to the source control points after the deformation process. By setting the source control points and target control points of the image region to be processed, the approximate movement trajectory of the target object in the image region to be processed can be planned, and the jumping effect of the target object can be realized, thereby enhancing the attractiveness of the target object in the image.

[0058] In one embodiment, the source control points and target control points of the image region to be processed can be set by the operating object. That is to say, the operating object determines the source control points of the image region to be processed and the target control points corresponding to the source control points on the terminal device, so that the target object in the image region to be processed can be transformed to the position corresponding to the target control points through the deformation process.

[0059] In one embodiment, the number of source control points of the image region to be processed can be multiple (for example, 5), and the number of target control points corresponding to the source control points can also be multiple (for example, 5). The number of source control points and the number of target control points are the same. In essence, it is the change in the position of the control points at different times.

[0060] S203. Perform a deformation process on the image region to be processed according to the source control points and target control points to obtain a deformed region image, and the target object is displayed as a second form in the deformed region image.

[0061] In the embodiments of the present application, by determining the source control points and target control points, the terminal device can perform a deformation process on the image region to be processed corresponding to the source control points to the target control points to obtain a deformed region image. At this time, the target object is displayed as a second form in the deformed region image. By generating a deformed region image including the second form of the target object, it is convenient for the target object to switch between the first form and the second form.

[0062] S204. Generate a target video based on the original image and the deformed area image, where the target object in the target video dynamically switches between a first form and a second form.

[0063] In the embodiment of the present application, the target object is in the first form in the original image and in the second form in the deformed area image. The terminal device can generate a target video based on the original image and the deformed area image, and in the target video, the target object is dynamically displayed between the first form and the second form. By generating an animated effect video, the attractiveness of the target object in the original image can be effectively improved, attracting the attention of the viewing object, thereby bringing a better advertising conversion effect.

[0064] In one embodiment, the method for dynamically displaying the target object in the target video can be to dynamically switch between the first form and the second form (for example, the terminal device generates two images respectively including the first form and the second form of the target object, and then uses the two images as two frames of the video and plays them repeatedly to achieve the purpose of dynamic display); the method for dynamically displaying the target object in the target video can also be to generate an animated effect video through video processing technology that enables the target object to have a jumping effect between the first form and the second form in the video.

[0065] In one embodiment, the target object includes keywords. That is, when performing animated effect processing on the keywords in the original image, the extraction of the image of the area to be processed from the original image can be achieved according to the following steps.

[0066] (a1). Perform text recognition processing on the foreground content of the original image to obtain candidate texts.

[0067] In the embodiment of the present application, the original image can include two categories: foreground content and background content. The foreground content can include people, goods, text, etc. (mainly close-up content); the background content can include the environment, background color, etc. (mainly long-distance content, or virtual content that sets off the foreground content). The candidate texts are all the texts that can be recognized in the original image.

[0068] In one embodiment, the terminal device can train the advertising image through a segmentation model (such as the Pyramid Scene Parsing Network, PSPNet model), and the data set can use 20,000 images to achieve better model processing accuracy. When performing text recognition processing, the optical character recognition technology (Optical Character Recognition, OCR) can be used, and the open-source PaddleOCR tool library can be used to detect the text area and recognize the text content to obtain candidate texts.

[0069] (a2) Perform a matching process on the candidate text with the keyword library to determine the matching keywords in the candidate text.

[0070] In the embodiments of the present application, the terminal device can perform a matching process on the candidate text with the keyword library to determine the matching keywords in the candidate text. The keyword library can be a text content, advertising core selling points, etc. Only the words that match the keyword library are considered matching keywords. The keyword library can be pre-set by the administrator, so as to achieve a better keyword matching accuracy, and then improve the accuracy of extracting the image of the area to be processed.

[0071] In one embodiment, the matching process can adopt steps of word segmentation, word embedding (word embedding is to convert words into numerical data, and the algorithm can adopt sent2vec), and vector distance calculation method (the threshold of the vector is 0.7) to implement the matching process between the candidate text and the keyword library.

[0072] (a3) Determine the image area corresponding to the matching keyword in the original image as the image of the area to be processed.

[0073] Through the above method, the terminal device can automatically obtain the matching keywords that match the keyword library, and then perform deformation processing on the image area corresponding to the matching keywords, avoiding invalid processing of non-critical areas, reducing the computational amount of the deformation processing, and thus improving the efficiency of the deformation processing.

[0074] Please refer to Figure 3A , this figure is a schematic diagram of the effect of an image processing method proposed in the present application. The left figure is the original image (such as an image of a notice type), including the target object 301 (such as "This website will stop operating in April 2022. We apologize for any inconvenience caused to you") and the date (such as "January 2022"). Through the image processing method proposed in the present application, the target video corresponding to the original image (such as the text jumping video file 1) can be generated. In the target video, the target object can be dynamically displayed before the first form and the second form.

[0075] Please refer to Figure 3B , this figure is a schematic diagram of the effect of an image processing method proposed in the present application. The left figure is the original image (such as an image of an advertisement type), including the target object 302 (such as "Click to listen") and the title (such as "In-car radio"). Through the image processing method proposed in the present application, the target video corresponding to the original image (such as the text jumping video file 2) can be generated. In the target video, the target object can be dynamically displayed before the first form, the second form, the third form, etc. (such as switching and displaying between different forms).

[0076] Based on the above image processing method, this application constructs a tool for generating videos from pictures. Through this tool, a target video with dynamic effects can be produced for the original image. Please refer to Figure 3C , which is a flowchart of an image processing method proposed in this application. By inputting the original image into the tool for generating videos from pictures, and the operator selects the type of dynamic effect (for example, determining the source control points and target control points in the original image), the tool for generating videos from pictures will automatically identify the target object (such as advertising copy), generate dynamic effects for the target object, and finally output the target video.

[0077] When processing an advertising image (the target object in the advertising image includes keywords), that is, when deforming the keywords in the original image, in the process of generating the above-mentioned video from pictures, it can specifically include steps such as text recognition processing and keyword matching processing. Specifically, please refer to Figure 3D , which is a schematic flowchart of the keyword image processing method proposed in this application. The terminal device first performs advertising foreground area recognition, text detection and recognition on the advertising picture to obtain candidate texts; then the terminal device performs key copy recognition on the candidate texts (that is, performs keyword library matching processing) to obtain matching keywords; the terminal device then performs key copy erasing processing on the advertising image, generates a text jumping effect according to the image of the area to be processed corresponding to the matching keyword, and generates an advertising video according to the text jumping effect and the advertising picture after key copy erasing processing.

[0078] Through the above tool for generating videos from pictures, the text on the original image is erased and then reproduced with the special effect of text jumping, which can automatically generate a dynamic effect video, improve the production efficiency of the dynamic effect video, reduce the production cost, expand the scale of dynamic effect materials, and through the above method, the produced materials can also be mass-produced to adapt to a huge number and various types of business scenarios, improving the applicability of dynamic effect production.

[0079] In an embodiment, when the terminal device performs dynamic effect processing on the keywords in the original image, the above-mentioned source control points for determining the image of the area to be processed can be realized according to the following steps.

[0080] (b1) Obtain the length feature of the image of the area to be processed, and the length feature includes at least one of the image length of the image of the area to be processed and the number of characters in the image of the area to be processed.

[0081] (b2) Determine at least one source control point of the image of the area to be processed according to the length feature, and the number of source control points is proportional to the length feature of the image of the area to be processed.

[0082] In the embodiments of the present application, in order to align the text with the source control points, the present application considers the relationship between the number of text characters and the number of source control points. Through experiments, it is found that when the number of source control points is too large, the movement amplitude of the text will become smaller, thus affecting the visual perception; when the number of source control points is too small, the area to be processed will be overly distorted, resulting in unclear text. Finally, the present application sets a threshold range for the number of control points based on the length feature of the image in the area to be processed. Within the threshold range, the number of source control points is positively correlated with the length feature. Specifically, the length feature of the image in the area to be processed can be features such as the image length of the area to be processed or the number of text characters in the area to be processed. When the length feature is larger, the number of source control points in the area to be processed is also larger. Through the above method, the visual perception of the animation effect can be improved while ensuring the clarity of the target video.

[0083] Please refer to Figure 3E , which is a schematic diagram of determining the source control points of the image in the area to be processed proposed by the present application. The target object in the image in the area to be processed is the keyword "image processing", and the number of text characters is 4. At this time, a source control point can be set respectively above and below each text character (the black dots in the figure are the source control points). Finally, 8 source control points are obtained to control the deformation of the image in the area to be processed. It should be noted that the number and position of the above-set source control points are only exemplary. When actually setting the source control points, they should be flexibly set according to the specific business situation to achieve a better animation effect perception.

[0084] The beneficial effects of the present application are as follows:

[0085] The present application extracts the image in the area to be processed including the target object from the original image for deforming the target object, avoiding processing the entire image, reducing the calculation amount, and thus improving the efficiency of the deformation processing; by determining the source control points of the image in the area to be processed and their corresponding target control points, the terminal device can plan the approximate movement trajectory of the target object in the image in the area to be processed; the terminal device can perform the deformation processing corresponding to the source control points to the target control points on the image in the area to be processed to obtain a deformed area image, and then generate a target video based on the original image and the deformed area image, so as to realize the dynamic display of the target object between the first form and the second form, which can effectively enhance the attractiveness of the target object in the original image, attract the attention of the viewing object, and thus bring a better advertising conversion effect. Compared with making an animation video manually, the present application can automatically generate an animation video based on the original graphics, improve the production efficiency of the animation video, and reduce the production cost.

[0086] The present application also proposes to perform dynamic effect processing on the overall image area corresponding to multiple characters, avoiding processing each character individually, improving the overall coordination of the target object before and after deformation, and enhancing the viewing experience of the object. The present application also proposes that the candidate text can be matched with keyword libraries such as the text content word library and the core selling point word library of advertisements to determine the matching keywords in the candidate text, thereby achieving a higher keyword matching accuracy and further improving the accuracy of extracting the image of the area to be processed. Moreover, the above method only performs deformation processing on the image area corresponding to the matching keywords, avoiding ineffective processing of non-critical areas, reducing the computational amount of the deformation processing, and thus improving the efficiency of the deformation processing. The present application also constructs a tool for generating videos from pictures, which automatically generates dynamic effect videos, improves the production efficiency of dynamic effect videos, reduces the production cost, expands the scale of dynamic effect materials, and can also perform large-scale production on the prepared materials through the above method to adapt to a huge number and various types of business scenarios, improving the applicability of dynamic effect production. The present application also proposes to set a threshold range for the number of control points based on length features such as the image length and the number of characters of the image to be processed. Within the threshold range, the number of source control points is positively correlated with the length features. This improves the visual perception of the dynamic effect while ensuring the clarity of the target video.

[0087] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of an image processing method provided by an exemplary embodiment of the present application. Taking the application of this method to a terminal device (referring to the terminal device 101 in Figure 1 ) as an example, the method may include the following steps:

[0088] S401. Extract the image of the area to be processed from the original image. The image of the area to be processed includes a target object, and the target object is displayed in a first form in the image of the area to be processed.

[0089] In one embodiment, extracting the image of the area to be processed from the original image may be that the terminal device cuts out the image of the area to be processed from the original image and obtains an image that does not include the image of the area to be processed. Subsequently, the terminal device can perform image restoration processing based on the image that does not include the image of the area to be processed, and then perform the production of the target video based on the image after the image restoration processing. Extracting the image of the area to be processed from the original image may also be that the terminal device extracts the image of the area to be processed from the original image and leaves the original image. Subsequently, the terminal device can perform an erasing operation on the target object based on the original image, then perform image restoration processing on the image after the erasing operation, and then perform the production of the target video based on the image after the image restoration processing. Through the above method, it is convenient to select a suitable method for image processing in different business environments, improving the flexibility of image processing.

[0090] In one embodiment, when the target object is a keyword, the above erasure operation is a copywriting erasure operation. After the copywriting erasure operation, an image restoration algorithm can be used for image restoration, such as the method provided by Resolution-robust LargeMask Inpainting with Fourier Convolutions. By constructing an advertisement dataset (for example, 100,000 images), retrain to adapt to the images in the advertisement scenario.

[0091] S402. Determine the source control points of the image in the area to be processed, and determine the target control points corresponding to the source control points.

[0092] Among them, for the specific implementation manners of steps S401 to S402, refer to the relevant descriptions of steps S201 to S202 in the foregoing embodiments, which will not be elaborated here.

[0093] S403. Determine the source following points corresponding to the source control points. The source following points are pixel points in the image in the area to be processed that satisfy the first distance condition from the source control points.

[0094] In the embodiments of the present application, the source following points are pixel points near the source control points in the image in the area to be processed, and the number of source following points corresponding to one source control point is at least one (generally multiple). By determining the source control points and the source following points, the spatial positions of all pixel points in the original image can be determined. It should be noted that if it is necessary to determine the spatial positions of all pixel points in the original image, generally multiple source control points need to be set, and each source control point corresponds to multiple source following points near this control point. Through the above method, using multiple source control points and the multiple source following points corresponding to each source control point among the multiple source control points, jointly determine the spatial positions of all pixel points in the original image.

[0095] S404. Determine the target following points corresponding to the source following points.

[0096] In the embodiments of the present application, the target following points are obtained by deforming the source following points. The number of source following points and target following points is the same. In essence, it is the change in the positions of the following points at different times. Based on the above method, the source control points, target control points, and source following points have been determined. When the terminal device further determines the target following points corresponding to the source following points, the deformed area image after the deformation processing of the image in the area to be processed can be obtained.

[0097] In one embodiment, the above determination of the target following points corresponding to the source following points can be implemented according to the following steps.

[0098] (c1). Determine the fitting curve coefficients of the source control points, and determine the fitting function corresponding to the source control points according to the fitting curve coefficients.

[0099] In the embodiments of the present application, the terminal device needs to establish a fitting function (i.e., a fitting curve function) near each source control point. Each source control point in the image of the area to be processed corresponds to a set of fitting curve coefficients for defining the shape of the fitting curve near the source control point. Therefore, when calculating the fitting function corresponding to the source control point, only the fitting curve coefficients of the source control point need to be calculated.

[0100] In one embodiment, the above-mentioned determination of the fitting curve coefficients of the source control point can be implemented according to the following steps.

[0101] (c11) Determine a plurality of sampling points corresponding to the source control point. The sampling points are pixel points in the image of the area to be processed that satisfy the second distance condition with respect to the source control point. Among them, the second distance condition may be the same as or different from the first distance condition.

[0102] (c12) Determine the fitting curve coefficients of the source control point according to the plurality of sampling points and the weights corresponding to the plurality of sampling points; among them, the weight corresponding to the sampling point is proportional to the distance between the sampling point and the source control point.

[0103] In the embodiments of the present application, the value of the fitting curve coefficients of each source control point only considers its adjacent sampling points. Moreover, the closer the sampling point is to the source control point, the greater the contribution, and the farther the sampling point is from the source control point, the smaller the contribution. For sampling points that are far from the source control point (i.e., the distance between the sampling point and the source control point does not satisfy the second distance condition), they are not considered. By adjusting the coefficients, the weighted sum of squares of the difference between the values of the sampling points near the source control point and the values of the fitting function at the sampling points is minimized, and thus the optimization model J can be established.

[0104] (c2) Determine the target following point corresponding to the source following point according to the fitting function corresponding to the source following point and the source control point.

[0105] In the embodiments of the present application, the terminal device can determine the target following point corresponding to the source following point through the fitting function corresponding to the source control point and the source following point, that is, the new pixel point after the position of the source following point changes after the deformation process. Furthermore, the deformed area image after the deformation process of the image of the area to be processed can be obtained through the target control point and the target following point.

[0106] The deformation of text can be abstracted as the grid deformation of the image within the font bounding box. Please refer to Figure 5 , this figure is a schematic diagram of the deformation of the image of the area to be processed. The terminal device first grids the image, and then uses the source control points (7 control points in the left figure) to change to the target control points (7 control points in the right figure), thereby realizing the change of the entire image grid.

[0107] The above steps (c1)-(c2) are image deformation implemented based on the moving least squares method. Briefly, its principle is to control the deformation through some source control points and target control points. For each source control point to be deformed, the terminal device first performs deformation processing according to a preset deformation type (for example, affine transformation, similarity transformation, rigid transformation, etc.); then estimates a local coordinate transformation matrix by solving a least squares optimization objective function, and then applies the coordinate transformation matrix to the source control point to calculate the position of the point after deformation (in this solution, the positions of the source control points and target control points can be set by the operating object, that is, the source control points and target control points are known, and the coordinate transformation matrix can be deduced in reverse). The sequence composed of these source control points and target control points serves as the trajectory of the text movement, thereby generating the deformed area image.

[0108] The moving least squares method used in this application will be described in detail below.

[0109] The moving least squares method (MLS, Moving Least Squares) is an ideal method for fitting curves with a large number of discrete data. When the distribution of a large number of discrete data is relatively chaotic, using the traditional least squares method often requires piecewise fitting of the data, and in addition, it is necessary to avoid the problem that the fitting curves on adjacent segments are discontinuous and non-smooth. The MLS method does not require these cumbersome steps when dealing with the same problem, and is simple and easy to implement.

[0110] Compared with the traditional least squares method, the fitting function in the MLS method is not a polynomial, but a set of coefficient vector functions a j (x) and basis functions P j (x), where x is the spatial coordinate. The fitting function near a certain node node (that is, the source control point in this application) is u node (x), and the specific definition is as follows:

[0111]

[0112] Among them, x node is the spatial coordinate of the node, x is the coordinate of a certain position near the node node, a j is a set of coefficients used to define the fitting curve near the node node (that is, the fitting curve coefficients in this application), and P j (x) is a set of basis functions.

[0113] In the MLS method, it is necessary to adjust the coefficient a j to make the weighted sum of squares of the difference between the sampled point values near the node node and the sampled point values of the fitting function the smallest, thereby establishing an optimization model J. The form of the optimization model J is as follows:

[0114]

[0115]

[0116] Among them, x node is the spatial position of node node, P is the sampling point, and x p is the spatial position of sampling point p, and u p is the value of sampling point p. Here

[0117] In the above formula, w is the weight function, which ensures that the closer the sampling point is to node node, the greater its contribution to the optimization model J; while for the sampling points farther away from node node, the smaller their contribution to the optimization model J; for the sampling points whose distance from node node exceeds the distance condition, they will not affect the optimization model J. The function w can take the following form:

[0118]

[0119] Among them, when the distance between the sampling point P and node node is greater than a certain value (here it is 2), the value of the weight function w is zero. (In the above figure, |x| is the modulus of the vector x). Of course, the neighborhood threshold of this function can be set to an appropriate value as needed. Through the above method, the coefficients of all nodes in the problem domain can be finally solved, and the fitting curve over the entire domain can be obtained.

[0120] The above steps have obtained the optimization model J with coefficients When the value of J is the smallest, the fitting curve u node (x) near node node can be obtained. By taking the derivative of the optimization model, the following equation is obtained:

[0121]

[0122] When the value of the above derivative formula is 0, the optimization model J can determine the minimum value, and finally can be sorted into the following linear equations (also called normal equations):

[0123]

[0124] This linear equation system can be written in matrix form: Among them, M is the coefficient matrix composed of basis functions, is the unknown vector, that is, the coefficient on node node Therefore, the unknown vector can also be expressed as:

[0125] By solving the above linear equations, the coefficients at node can be obtained. Through the coefficients and the basis vector p j (x), a fitting function near node can be established. At this time, the form of the fitting function at any position is as follows:

[0126]

[0127] Substitute into it, and the following formula can be obtained:

[0128]

[0129] where

[0130] Through the above method, the fitting function and its corresponding value can be calculated at any point x, and the values of two adjacent points in space are very smooth.

[0131] S405. According to the source control points, source follower points, target control points and target follower points, perform deformation processing on the image of the area to be processed to obtain a deformed area image, and the target object is displayed as the second form in the deformed area image.

[0132] In the embodiment of the present application, according to the source control points and target control points, the image of the target object before change (i.e., the image of the area to be processed) can be determined; according to the target control points and target follower points, the image of the target object after change (i.e., the deformed area image) can be determined. In the image of the area to be processed, the target object is displayed as the first form, and in the deformed area image, the target object is displayed as the second form.

[0133] S406. Generate a target video according to the original image and the deformed area image, and the target object in the target video dynamically switches between the first form and the second form.

[0134] In one embodiment, the process of generating the target video according to the original image and the deformed area image can be implemented according to the following steps: Take the original image as the first video frame; fuse the deformed area image and the original image as the second video frame, and generate a continuous video according to the first video frame and the second video frame. The first video frame and the second video frame are cyclically displayed in this video so that the target object dynamically switches between the first form and the second form.

[0135] In one embodiment, in order to make the dynamic switching of the target object in the target video smoother, the above method may further include the following steps.

[0136] (d1) Determine intermediate control points based on source control points and target control points, and determine intermediate following points that match the intermediate control points based on source following points and target following points.

[0137] (d2) Perform deformation processing on the image of the area to be processed based on the source control points, source following points, intermediate control points, and intermediate following points to obtain an intermediate deformed image; the target object is displayed as a third form in the intermediate deformed image.

[0138] In the embodiments of the present application, the intermediate control points and intermediate following points are used to generate an intermediate deformed image, so that the dynamic display of the target object in the target video is smoother and the viewing experience is improved. Among them, the number of intermediate control points can be multiple, and the number of intermediate following points corresponding to the multiple intermediate control points can also be multiple. Each intermediate control point and its corresponding intermediate following point can form an intermediate deformed image.

[0139] In one embodiment, the above-mentioned extraction of the image of the area to be processed from the original image can be achieved according to the following steps: determine the area information of the image of the area to be processed from the original image, and perform image segmentation processing on the original image according to the area information to obtain the image of the area to be processed and the original image after removing the image of the area to be processed.

[0140] Among them, the area information of the image of the area to be processed is used to determine the image area to be processed. For example, the area information of the image of the area to be processed can be the contour information of the image of the area to be processed, etc. Through the contour information of the image of the area to be processed, the original image can be segmented into the image of the area to be processed and the original image after removing the image of the area to be processed. At this time, the terminal device can perform image restoration processing based on the original image after removing the image of the area to be processed, and then produce a target video based on the image after the image restoration processing.

[0141] In one embodiment, the above-mentioned determination of intermediate control points based on source control points and target control points, and the determination of intermediate following points that match the intermediate control points based on source following points and target following points can be obtained through difference processing (such as linear interpolation processing). Taking the source control points and target control points as an example, if the spatial coordinates of the source control points are (1, 1, 1) and the spatial coordinates of the target control points are (7, 7, 7), then through linear interpolation processing, two intermediate control points (3, 3, 3) and (5, 5, 5) can be obtained, so that the source control points can change from (1, 1, 1) to the target control points (7, 7, 7) slowly and smoothly through (3, 3, 3) and (5, 5, 5).

[0142] Based on the above steps (d1)-(d2), the generation of the target video from the original image and the deformed region image can be achieved according to the following steps: generating the target video from the original image, the intermediate deformed image, and the deformed region image, where the target object in the target video dynamically switches between the first form, the third form, and the second form.

[0143] In one embodiment, the process of generating the target video from the original image and the deformed region image can be achieved according to the following steps: taking the original image as the first video frame; fusing the intermediate deformed image and the original image as the intermediate video frame (which can include multiple frames); fusing the deformed region image and the original image as the second video frame, and generating a continuous video based on the first video frame, the intermediate video frame, and the second video frame, where the first video frame, the intermediate video frame, and the second video frame are cyclically displayed in this video, so that the target object dynamically switches between the first form, the third form (here the third form corresponds to multiple intermediate video frames, which can include multiple forms), and the second form.

[0144] In one embodiment, the generation of the target video from the original image, the intermediate deformed image, and the deformed region image can be achieved according to the following steps.

[0145] (e1) Perform image restoration processing on the original image after removing the image of the area to be processed to obtain a restored image.

[0146] In one embodiment, when the target object is a keyword, the above erasing operation is a copywriting erasing operation. After the copywriting erasing operation, an image restoration algorithm can be used for image restoration, such as the method provided by Resolution-robust LargeMask Inpainting with Fourier Convolutions. By constructing an advertising dataset (for example, 100,000 images), retrain to adapt to the images in the advertising scenario.

[0147] (e2) Fuse the restored image and the intermediate deformed image to obtain an intermediate video frame image.

[0148] (e3) Fuse the restored image and the deformed region image to obtain a target video frame image.

[0149] (e4) Take the original image as the initial video frame image, and generate the target video based on the initial video frame image, the intermediate video frame image, and the target video frame image.

[0150] After repairing the image and then fusing it with the intermediate deformed image and the deformed area image, a video frame image is obtained, and then a target video is generated, which can ensure the integrity of the foreground content and the background content in the area where the target object is located in the video, guarantee the quality of the video, and further improve the viewing experience.

[0151] This application provides a method for deforming a target object from an image in a region to be processed into a deformed region image, and a method for generating an intermediate deformed image between the image in the region to be processed and the deformed region image, so as to achieve dynamic display of the target object in different forms in the target video. In an embodiment, the terminal device can continue to obtain additional control points input by the operating object (i.e., control points set by the operating object for the desired changes of the target object). According to the method proposed in this application, the terminal device can determine the additional following points corresponding to the additional control points, so as to determine the additional deformed images (there can be multiple additional following points, and there are also multiple corresponding additional deformed images, and each additional deformed image is calculated based on the previous additional deformed image). The terminal device can generate a target video with multiple forms through the original image, the deformed region image, the additional deformed images, and the intermediate deformed images between adjacent two images.

[0152] In an embodiment, the operating object can input multiple control points corresponding to, for example, the laws of human motion on the terminal device, so that the target object can simulate the laws of human motion in the video, making the jumping effect of the target object more realistic and coordinated, effectively enhancing the attractiveness of the key core text information in the video generated from pictures, presenting a special effect of text jumping in the form of text expression, greatly improving the effect of the video, attracting the attention of the object, and thus bringing a better advertising conversion effect.

[0153] In an embodiment, the target control points corresponding to the source control points, the intermediate control points corresponding to the source control points, and the additional control points corresponding to the source control points can be automatically generated by the terminal device. Specifically, multiple types of jumping types are preset in the terminal device (for example, jumping clockwise at the three vertices of a triangle, or twisting in imitation of the dancing rhythm of a human body, etc.). When the terminal device obtains the original image and the source control points, it can generate the corresponding target video according to the source control points and the target jumping type in the preset jumping types, enabling the automatic generation of the target video and greatly improving the generation efficiency of the target video.

[0154] Please refer to Figure 5, this figure is a schematic diagram of the deformation of an image of an area to be processed provided by an embodiment of the present application. The figure includes an image 501 of the area to be processed and a deformed area image 502. The terminal device first sets source control points for the image 501 of the area to be processed (i.e., the black dots in 501, and the positions of the source control points can be determined by the input information of the operation object by the terminal device); then the terminal device performs grid processing on the image, and then obtains the deformed area image 502 through the deformation method proposed by the present application. In the deformed area image 502, the positions of the source control points have changed (that is, they have changed from the positions where the source control points are located to the positions where the target control points are located).

[0155] The present application can be applied to a picture-to-video tool in the advertising field, which is used to perform dynamic effect processing on all advertising materials containing text to increase the interest of the text. The present application has the following effects: The first point is to increase the proportion of videos in the advertising materials and enrich the advertising material library. The second point is that the already produced materials can be processed again, and the dynamic effect processing is fully automatic, greatly improving the reuse rate of the materials. The third point is that a more realistic text jumping dynamic effect can be obtained, which can enhance the attention of the viewing object by enhancing the text, making the viewing object more impressed by the information that the advertisement core wants to convey, and improving the effect of the advertising materials.

[0156] It should be noted that various models proposed by the present application can be optimized accordingly, and the same effect can be achieved by using similar models or methods (such as text OCR recognition, foreground area recognition, recognition rules for key text, text image restoration, text special effect generation schemes, etc.), and the present application does not limit them. The present application mainly takes advertising pictures as an example, and through the production of dynamic effects for the advertising keywords in the advertising pictures, the effect of the advertising keywords is improved. In addition, the method proposed by the present application can also be applied to many fields such as short videos, reading, design, etc. For example, the static titles in short videos are dynamically displayed, and the load-bearing wall parts in house design images are dynamically displayed, etc. The present application does not limit the application fields. The idea of the text dynamic effect proposed by the present application can be extended to design more movement trajectories to make the text move, and the fine-tuning of such action sequences can be realized based on the method proposed by the present application.

[0157] The beneficial effects of the present application are as follows:

[0158] In the process of obtaining the deformed area image in this application, at least one source control point and a source following point corresponding to each source control point that meets the first distance condition are determined, so as to jointly determine the spatial positions of all pixel points in the original image, that is, the deformed area image is obtained. This application proposes to implement image deformation based on the moving least squares method, which improves the accuracy and efficiency of image deformation. This application also proposes to determine an intermediate deformed image through intermediate control points and intermediate following points, so that the dynamic display of the target object in the target video is smoother and the viewing experience is improved. This application also proposes to fuse the image after repair with the intermediate deformed image and the deformed area image to obtain a video frame image, and then generate a target video, which can ensure the integrity of the foreground content and background content in the area where the target object is located in the video, ensure the quality of the video, and further improve the viewing experience. This application also proposes that the terminal device can automatically generate multiple control points corresponding to, for example, the human motion law, so that the target object can simulate the human motion law in the video, making the jumping effect of the target object more realistic and coordinated, effectively enhancing the attractiveness of the key core text information in the video generated by pictures, presenting a special effect of text jumping in the text presentation form, greatly enhancing the effect of the video, attracting the attention of the object, and thus bringing a better advertising conversion effect.

[0159] Please refer to Figure 6 , Figure 6 which is a schematic block diagram of an image processing device provided by an embodiment of this application. Specifically, the image processing device may include:

[0160] An acquisition module 601, configured to extract an image of a to-be-processed area from the original image, where the image of the to-be-processed area includes a target object, and the target object is displayed in a first form in the image of the to-be-processed area;

[0161] A processing module 601, configured to determine source control points of the image of the to-be-processed area and determine target control points corresponding to the source control points;

[0162] The processing module 601 is further configured to perform deformation processing on the image of the to-be-processed area according to the source control points and the target control points to obtain a deformed area image; the target object is displayed in a second form in the deformed area image;

[0163] A generation module 602, configured to generate a target video according to the original image and the deformed area image, and the target object in the target video dynamically switches between the first form and the second form.

[0164] Optionally, when the processing module 602 is configured to perform deformation processing on the image of the to-be-processed area according to the source control points and the target control points to obtain a deformed area image, it is specifically configured to:

[0165] Determine the source following points corresponding to the above source control points, where the above source following points are pixel points in the image of the area to be processed that satisfy the first distance condition with respect to the above source control points;

[0166] Determine the target following points corresponding to the above source following points;

[0167] Perform a deformation process on the image of the area to be processed according to the above source control points, the above source following points, the above target control points, and the above target following points to obtain a deformed area image.

[0168] Optionally, when the processing module 602 is used to determine the target following points corresponding to the above source following points, it is specifically used for:

[0169] Determine the fitting curve coefficients of the above source control points, and determine the fitting function corresponding to the above source control points according to the above fitting curve coefficients;

[0170] Determine the target following points corresponding to the above source following points according to the above source following points and the fitting function corresponding to the above source control points.

[0171] Optionally, when the processing module 602 is used to determine the fitting curve coefficients of the above source control points, it is specifically used for:

[0172] Determine multiple sampling points corresponding to the above source control points, where the sampling points are pixel points in the image of the area to be processed that satisfy the second distance condition with respect to the above source control points;

[0173] Determine the fitting curve coefficients of the above source control points according to the above multiple sampling points and the weights corresponding to the above multiple sampling points; wherein, the weight corresponding to a sampling point is proportional to the distance between the sampling point and the above source control point.

[0174] Optionally, the above processing module 602 is further used for:

[0175] Determine intermediate control points according to the above source control points and the above target control points, and determine intermediate following points that match the above intermediate control points according to the above source following points and the above target following points;

[0176] Perform a deformation process on the image of the area to be processed according to the above source control points, the above source following points, the above intermediate control points, and the above intermediate following points to obtain an intermediate deformed image; the above target object is displayed in a third form in the above intermediate deformed image.

[0177] When the above generating module 603 is used to generate a target video according to the above original image and the above deformed area image, it is specifically used for:

[0178] Generate a target video based on the above-mentioned original image, the above-mentioned intermediate deformed image, and the above-mentioned deformed region image. The target object in the above-mentioned target video dynamically switches between the above-mentioned first form, the above-mentioned third form, and the above-mentioned second form.

[0179] Optionally, when the processing module 602 is used to extract the image of the area to be processed from the original image, it is specifically used for:

[0180] Determine the region information of the image of the area to be processed from the original image, perform image segmentation processing on the above-mentioned original image according to the above-mentioned region information, obtain the above-mentioned image of the area to be processed, and obtain the original image after removing the above-mentioned image of the area to be processed;

[0181] When the processing module 602 is used to generate a target video based on the above-mentioned original image, the above-mentioned intermediate deformed image, and the above-mentioned deformed region image, it is specifically used for:

[0182] Perform image restoration processing on the original image after removing the above-mentioned image of the area to be processed to obtain a restored image;

[0183] Fuse the above-mentioned restored image and the above-mentioned intermediate deformed image to obtain an intermediate video frame image;

[0184] Fuse the above-mentioned restored image and the above-mentioned deformed region image to obtain a target video frame image;

[0185] Use the above-mentioned original image as the initial video frame image, and generate a target video based on the above-mentioned initial video frame image, the intermediate video frame image, and the above-mentioned target video frame image.

[0186] Optionally, the above-mentioned target object includes keywords. When the acquisition module 601 is used to extract the image of the area to be processed from the original image, it is specifically used for:

[0187] Perform text recognition processing on the foreground content of the original image to obtain candidate texts;

[0188] Match the above-mentioned candidate texts with a keyword library to determine the matching keywords in the above-mentioned candidate texts;

[0189] Determine the image region corresponding to the above-mentioned matching keywords in the above-mentioned original image as the image of the area to be processed.

[0190] It should be noted that the functions of the functional modules of the image processing device in the embodiments of the present application can be specifically implemented according to the methods in the above-mentioned method embodiments. The specific implementation process can refer to the relevant descriptions in the above-mentioned method embodiments and will not be elaborated here.

[0191] Please refer to Figure 7 , Figure 7It is a schematic block diagram of a computer device provided by an embodiment of the present application. The intelligent terminal in this embodiment shown in the figure may include: a processor 701, a storage device 702, and a communication interface 703. Data interaction can be carried out among the above-mentioned processor 701, storage device 702, and communication interface 703.

[0192] The above storage device 702 may include a volatile memory, such as a random-access memory (RAM); the storage device 702 may also include a non-volatile memory, such as a flash memory, a solid-state drive (SSD), etc.; the above storage device 702 may further include a combination of the above types of memories.

[0193] The above processor 701 may be a central processing unit (CPU). In one embodiment, the above processor 701 may also be a Graphics Processing Unit (GPU). The above processor 701 may also be a combination of a CPU and a GPU. In one embodiment, the above storage device 702 is used to store program instructions, and the above processor 701 may call the above program instructions to perform the following operations:

[0194] Extract an image of a region to be processed from the original image, where the image of the region to be processed includes a target object, and the target object is displayed in a first form in the image of the region to be processed;

[0195] Determine the source control points of the image of the region to be processed, and determine the target control points corresponding to the source control points;

[0196] Perform a deformation process on the image of the region to be processed according to the source control points and the target control points to obtain a deformed region image; the target object is displayed in a second form in the deformed region image;

[0197] Generate a target video according to the original image and the deformed region image, and the target object in the target video dynamically switches between the first form and the second form.

[0198] Optionally, when the above processor 701 is used to perform a deformation process on the image of the region to be processed according to the source control points and the target control points to obtain a deformed region image, it is specifically used for:

[0199] Determine the source following points corresponding to the above source control points, where the above source following points are pixel points in the image of the area to be processed that satisfy the first distance condition with respect to the above source control points;

[0200] Determine the target following points corresponding to the above source following points;

[0201] Perform a deformation process on the image of the area to be processed according to the above source control points, the above source following points, the above target control points, and the above target following points to obtain a deformed area image.

[0202] Optionally, when the above processor 701 is used to determine the target following points corresponding to the above source following points, it is specifically used for:

[0203] Determine the fitting curve coefficients of the above source control points, and determine the fitting function corresponding to the above source control points according to the above fitting curve coefficients;

[0204] Determine the target following points corresponding to the above source following points according to the above source following points and the fitting function corresponding to the above source control points.

[0205] Optionally, when the above processor 701 is used to determine the fitting curve coefficients of the above source control points, it is specifically used for:

[0206] Determine multiple sampling points corresponding to the above source control points, where the sampling points are pixel points in the image of the area to be processed that satisfy the second distance condition with respect to the above source control points;

[0207] Determine the fitting curve coefficients of the above source control points according to the above multiple sampling points and the weights corresponding to the above multiple sampling points; wherein, the weight corresponding to a sampling point is proportional to the distance between the sampling point and the above source control point.

[0208] Optionally, the above processor 701 is further used for:

[0209] Determine intermediate control points according to the above source control points and the above target control points, and determine intermediate following points that match the above intermediate control points according to the above source following points and the above target following points;

[0210] Perform a deformation process on the image of the area to be processed according to the above source control points, the above source following points, the above intermediate control points, and the above intermediate following points to obtain an intermediate deformed image; the above target object is displayed in a third form in the above intermediate deformed image.

[0211] When the above processor 701 is used to generate a target video according to the above original image and the above deformed area image, it is specifically used for:

[0212] Generate a target video based on the above-mentioned original image, the above-mentioned intermediate deformed image, and the above-mentioned deformed region image. The target object in the above-mentioned target video dynamically switches between the above-mentioned first form, the above-mentioned third form, and the above-mentioned second form.

[0213] Optionally, when the above-mentioned processor 701 is used to extract the image of the area to be processed from the original image, it is specifically used for:

[0214] Determine the region information of the image of the area to be processed from the original image, perform image segmentation processing on the above-mentioned original image according to the above-mentioned region information, obtain the above-mentioned image of the area to be processed, and obtain the original image after removing the above-mentioned image of the area to be processed;

[0215] When the above-mentioned processor 701 is used to generate a target video based on the above-mentioned original image, the above-mentioned intermediate deformed image, and the above-mentioned deformed region image, it is specifically used for:

[0216] Perform image restoration processing on the original image after removing the above-mentioned image of the area to be processed to obtain a restored image;

[0217] Fuse the above-mentioned restored image and the above-mentioned intermediate deformed image to obtain an intermediate video frame image;

[0218] Fuse the above-mentioned restored image and the above-mentioned deformed region image to obtain a target video frame image;

[0219] Use the above-mentioned original image as the initial video frame image, and generate a target video based on the above-mentioned initial video frame image, the intermediate video frame image, and the above-mentioned target video frame image.

[0220] Optionally, the above-mentioned target object includes keywords. When the above-mentioned processor 701 is used to extract the image of the area to be processed from the original image, it is specifically used for:

[0221] Perform text recognition processing on the foreground content of the original image to obtain candidate texts;

[0222] Match the above-mentioned candidate texts with a keyword library to determine the matching keywords in the above-mentioned candidate texts;

[0223] Determine the image region corresponding to the above-mentioned matching keywords in the above-mentioned original image as the image of the area to be processed.

[0224] In specific implementation, the processor 701, the storage device 702, and the communication interface 703 described in the embodiments of the present application may execute the implementation manners described in the relevant embodiments of the image processing method provided by the embodiments of the present application Figure 2 or Figure 4 execute the implementation manners described in the relevant embodiments of the image processing method provided by the embodiments of the present application Figure 6The implementation methods described in the relevant embodiments of the provided image processing device will not be repeated here.

[0225] In the several embodiments provided in the present application, it should be understood that the disclosed methods, devices and systems can be implemented in other ways. For example, the device embodiments described above are merely schematic; for example, the division of the units is only a logical function division, and there may be other division methods in actual implementation; for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0226] In addition, it should be pointed out here that: the embodiment of the present application also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the image processing device mentioned above, and the computer program includes program instructions. When the processor executes the above program instructions, it can execute the above Figure 2 , Figure 4 The method in the corresponding embodiment, therefore, will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated. For technical details not disclosed in the computer-readable storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application. As an example, the program instructions can be deployed on a computer device, or executed on multiple computer devices located in one location, or, executed on multiple computer devices distributed in multiple locations and interconnected by a communication network, and multiple computer devices distributed in multiple locations and interconnected by a communication network can constitute a blockchain system.

[0227] According to one aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device can perform the above Figure 2 , Figure 4 The method in the corresponding embodiment will therefore not be described in detail here.

[0228] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The above program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above various methods. Among them, the above storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0229] The above-disclosed are only some embodiments of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of the above embodiments, and the equivalent changes made according to the claims of the present application still fall within the scope covered by the invention.

Claims

1. An image processing method, characterized in that, The method includes: Extracting an image of a region to be processed from an original image, where the image of the region to be processed includes a target object, and the target object is displayed in a first form in the image of the region to be processed; Determining source control points of the image of the region to be processed, and determining target control points corresponding to the source control points; Determining source following points corresponding to the source control points, where the source following points are pixel points in the image of the region to be processed that satisfy a first distance condition with respect to the source control points; Determining a plurality of sampling points corresponding to the source control points, where the sampling points are pixel points in the image of the region to be processed that satisfy a second distance condition with respect to the source control points; Determining fitting curve coefficients of the source control points according to the plurality of sampling points and weights corresponding to the plurality of sampling points; wherein, the weight corresponding to a sampling point is proportional to the distance between the sampling point and the source control point; Determining a fitting function corresponding to the source control points according to the fitting curve coefficients; Determining target following points corresponding to the source following points according to the fitting function corresponding to the source control points and the source following points; Performing a deformation process on the image of the region to be processed according to the source control points, the source following points, the target control points, and the target following points to obtain a deformed region image; the target object is displayed in a second form in the deformed region image; Generating a target video according to the original image and the deformed region image, where the target object in the target video dynamically switches between the first form and the second form.

2. The method according to claim 1, characterized in that The method further includes: Determining intermediate control points according to the source control points and the target control points, and determining intermediate following points matching the intermediate control points according to the source following points and the target following points; Performing a deformation process on the image of the region to be processed according to the source control points, the source following points, the intermediate control points, and the intermediate following points to obtain an intermediate deformed image; the target object is displayed in a third form in the intermediate deformed image; Wherein, the generating a target video according to the original image and the deformed region image includes: Generating a target video according to the original image, the intermediate deformed image, and the deformed region image, where the target object in the target video dynamically switches between the first form, the third form, and the second form.

3. The method according to claim 2, characterized in that, The extracting an image of a region to be processed from an original image includes: Determining region information of the image of the region to be processed from the original image, and performing image segmentation processing on the original image according to the region information to obtain the image of the region to be processed and the original image with the image of the region to be processed removed; Wherein, the generating a target video according to the original image, the intermediate deformed image, and the deformed region image includes: Performing image restoration processing on the original image with the image of the region to be processed removed to obtain a restored image; Fusing the restored image and the intermediate deformed image to obtain an intermediate video frame image; Fusing the restored image and the deformed region image to obtain a target video frame image; Taking the original image as the initial video frame image, a target video is generated according to the initial video frame image, the intermediate video frame image, and the target video frame image.

4. The method according to claim 1, wherein The target object includes keywords. Extracting the image of the area to be processed from the original image includes: Performing text recognition processing on the foreground content of the original image to obtain candidate texts; Matching the candidate texts with a keyword library to determine the matching keywords in the candidate texts; Determining the image area corresponding to the matching keywords in the original image as the image of the area to be processed.

5. An image processing apparatus, characterized in that, The device includes: An acquisition module, configured to extract the image of the area to be processed from the original image. The image of the area to be processed includes a target object, and the target object is displayed in a first form in the image of the area to be processed; A processing module, configured to determine the source control points of the image of the area to be processed and determine the target control points corresponding to the source control points; The processing module is further configured to: determine the source following points corresponding to the source control points. The source following points are pixel points in the image of the area to be processed that satisfy a first distance condition from the source control points; determine multiple sampling points corresponding to the source control points. The sampling points are pixel points in the image of the area to be processed that satisfy a second distance condition from the source control points; determine the fitting curve coefficients of the source control points according to the multiple sampling points and the weights corresponding to the multiple sampling points. The weight corresponding to a sampling point is proportional to the distance between the sampling point and the source control point; determine the fitting function corresponding to the source control points according to the fitting curve coefficients; determine the target following points corresponding to the source following points according to the fitting function corresponding to the source control points and the source following points; perform deformation processing on the image of the area to be processed according to the source control points, the source following points, the target control points, and the target following points to obtain a deformed area image. The target object is displayed in a second form in the deformed area image; A generation module, configured to generate a target video according to the original image and the deformed area image. The target object in the target video dynamically switches between the first form and the second form.

6. A computer device, characterized in that, Including: A processor, a storage device, and a communication interface. The processor, the communication interface, and the storage device are interconnected. Among them, the storage device stores executable program codes, and the processor is configured to call the executable program codes to implement the image processing method according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program. The computer program includes program instructions. The program instructions are executed by a processor to implement the image processing method according to any one of claims 1 to 4.

8. A computer program product, characterized in that, The computer program product includes a computer program or computer instructions. When the computer program or computer instructions are executed by a processor, they are used to implement the image processing method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Image processing method and device, storage medium and server

    CN109741277A

  • Image deformation method and device

    CN109903217A

  • Video generation method and device, electronic equipment and readable storage medium

    CN113961746A