Cartoon automatic playing method, device and equipment and computer storage medium

By merging and splitting the comic strip images and adjusting the playback speed according to the text information, the problem of image grid destruction in existing technologies is solved, resulting in a better user reading experience.

CN115577129BActive Publication Date: 2025-12-19HAPPINIS (BEIJING) CULTURE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110683429.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-21
Publication Date
2025-12-19
Estimated Expiration
2041-06-21

AI Technical Summary

Technical Problem

Existing autoplay technology can easily disrupt the integrity of image panels in strip comics, affecting the user's reading experience.

Method used

By acquiring the first image and text information from the comic strip, merging and splitting are performed based on their positional relationships to determine the image playback speed, and the playback speed is adjusted according to the text information.

Benefits of technology

Ensure image integrity during comic playback to improve the user's reading experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115577129B_ABST
    Figure CN115577129B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a kind of cartoon automatic playing method, device, equipment and computer storage medium.The cartoon automatic playing method, obtains the first image in target image and text information, and according to position relationship is merged, according to the image after merging target image is split, it is judged the density of text information in the image after splitting determines the automatic playing speed of target image, can be adapted to adjust the playing speed according to image text content while guaranteeing the complete image content of user reading, so that user obtains better reading experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of electronic information, and particularly relates to a method and device for automatically playing a comic, an equipment and a computer storage medium. BACKGROUND

[0002] With the popularity of electronic devices, people use electronic devices to read comics more and more. In order to provide users with a better reading experience, automatic playing technology has emerged. The current automatic playing technology mainly determines the number of texts in the current display area to determine the scrolling speed of the comic, that is, the more texts in the current display area, the slower the scrolling speed.

[0003] The above technical solution can realize automatic playing of a comic, but when the comic is in a strip form, the technical solution will destroy the complete image grid and affect the user's reading experience. SUMMARY

[0004] The embodiment of the present application provides a method and device for automatically playing a comic, an equipment and a computer storage medium, which can obtain first images and text information in a strip image and merge them according to their position relationship. The strip is split based on the merged image, the text information in the split image is obtained to determine the playing speed of the image, and the comic is played based on the playing speed, which can ensure that the user sees a relatively complete image during playing, and adaptively adjusts the playing speed according to the text information in the image to improve the user's reading experience.

[0005] In a first aspect, the embodiment of the present application provides a method for automatically playing a comic, which comprises:

[0006] Obtaining image data in a target image, the image data comprising first text information, a plurality of first images, and position relationship information of the first text information and the plurality of first images, the position relationship information comprising that the first text information overlaps with the first image, or the first text information and the plurality of first images partially overlap, or the first text information does not overlap with the plurality of first images;

[0007] Merging the first text information and the first image based on a preset rule and the position relationship information of the first text information and the first image to obtain a second image;

[0008] Splitting the target image based on the second image and a preset size to obtain a plurality of third images;

[0009] Obtaining image data of the third image, the image data comprising second text information, the second text information comprising texts and text density;

[0010] Determining the speed of automatic playing of the third image based on the second text information of the third image;

[0011] playing the third image based on the speed of the automatic playing.

[0012] According to one aspect of the present application, the speed of the automatic playing of the third image is determined based on the second text information of the third image, comprising:

[0013] determining a text weight according to the second text information;

[0014] determining the speed of the automatic playing of the third image according to the text weight and the character density.

[0015] In a second aspect, the embodiments of the present application provide a device for automatically playing a comic, which comprises:

[0016] a first obtaining module, configured to obtain image data in a target image, the image data comprising first text information, a plurality of first images, and position relationship information of the first text information and the plurality of first images, the position relationship information comprising that the first text information overlaps with the first image, or the first text information and the plurality of first images partially overlap, or the first text information does not overlap with the plurality of first images;

[0017] a merging module, configured to merge the first text information and the first image based on a preset rule and the position relationship information of the first text information and the first image to obtain a second image;

[0018] a splitting module, configured to split the target image based on the second image and a preset size to obtain a plurality of third images;

[0019] a second obtaining module, configured to obtain image data of the third image, the image data comprising second text information, the second text information comprising text and character density;

[0020] a determining module, configured to determine the speed of the automatic playing of the third image based on the second text information of the third image;

[0021] a playing module, configured to play the third image based on the speed of the automatic playing.

[0022] According to one aspect of the present application, the speed of the automatic playing of the third image is determined based on the second text information of the third image, comprising:

[0023] determining a text weight according to the second text information;

[0024] determining the speed of the automatic playing of the third image according to the text weight and the character density.

[0025] In a third aspect, the embodiments of the present application provide a device for automatically playing a comic, which comprises:

[0026] a processor, and a memory storing computer program instructions;

[0027] The processor reads and executes computer program instructions to implement the automatic playing method of the comic as any one of the first aspect.

[0028] In a fourth aspect, the embodiments of the present application provide a computer storage medium, having computer program instructions stored thereon, which, when executed by a processor, implement the automatic playing method of the comic as any one of the first aspect.

[0029] The automatic playing method of the comic, the device, the equipment and the computer storage medium provided by the embodiments of the present application can obtain the first image and the text information in the strip comic image, and merge the first image and the text information based on the position relationship between the first image and the text information. Meanwhile, the strip comic image can be split into appropriate lengths, and the playing speed of the image can be determined according to the text information in the split image. In this way, the playing speed can be adaptively adjusted according to the text density in the image when the complete image is ensured to be played, so that the user can obtain better reading experience. BRIEF DESCRIPTION OF DRAWINGS

[0030] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the embodiments of the present application will be briefly introduced. For those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0031] Figure 1 is a flowchart of the automatic playing method of the comic provided by the embodiments of the present application;

[0032] Figure 2 is a schematic diagram of the position relationship between the first image and the first text information provided by the embodiments of the present application;

[0033] Figure 3 is a schematic diagram of the position relationship between the first image and the first text information provided by the embodiments of the present application;

[0034] Figure 4 is a schematic diagram of the position relationship between the first image and the first text information provided by the embodiments of the present application;

[0035] Figure 5 is a schematic diagram of the obtained image data provided by the embodiments of the present application;

[0036] Figure 6 is a schematic diagram of the image data merging provided by the embodiments of the present application;

[0037] Figure 7 is a schematic diagram of the image data merging provided by the embodiments of the present application;

[0038] Figure 8 is a schematic diagram of the image data merging provided by the embodiments of the present application;

[0039] Figure 9 is a schematic diagram of image data merging provided by an embodiment of the present application;

[0040] Figure 10 is a schematic diagram of image data merging provided by an embodiment of the present application;

[0041] Figure 11 is a schematic diagram of image data merging provided by an embodiment of the present application;

[0042] Figure 12 is a schematic diagram of image splitting provided by an embodiment of the present application;

[0043] Figure 13 is a structural schematic diagram of a cartoon automatic playing device provided by an embodiment of the present application;

[0044] Figure 14 is a structural schematic diagram of a cartoon automatic playing device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0045] The features and exemplary embodiments of various aspects of the present application will be described below in detail, in order to make the purposes, technical solutions and advantages of the present application more clear and apparent, the present application will be further described in detail below in combination with the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, but not to limit the present application. The present application can be implemented without some of these specific details by those skilled in the art. The following description of the embodiments is only to provide a better understanding of the present application by showing examples of the present application.

[0046] It should be noted that, in this paper, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the elements defined by the statement "include" do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0047] At present, when the automatic playing of the cartoon is performed, the playing speed is determined according to all the texts and text density in the current page, the text information not belonging to the theme content of the cartoon is contained, and the integrity of the image cannot be guaranteed during the playing process.

[0048] To solve the problems in the prior art, the embodiment of the present application provides a method, device and equipment for automatically playing a comic and a computer storage medium. First, the automatic playing method provided by the embodiment of the present application is introduced.

[0049] Figure 1 A flowchart of the method for automatically playing a comic provided by an embodiment of the present application is shown. As shown in the figure, the method can include the following steps: Figure 1

[0050] S110, image data in a target image is acquired, the image data including first text information, a plurality of first images, and position relationship information of the first text information and the plurality of first images. The position relationship information includes that the first text information overlaps the first image, as shown in the figure; or the first text information and the plurality of first images partially overlap, as shown in the figure; or the first text information and the plurality of first images do not overlap, as shown in the figure. Figure 2 Figure 3 Figure 4

[0051] In some embodiments, the target image is a strip comic image, as shown in the figure. When the image data in the target image is acquired, the first text information and the plurality of first images are acquired by using an artificial intelligence (AI) model. Figure 5

[0052] In one example, the AI model can include two: a grid detection model and a text recognition model. The grid detection model is used to acquire the first image in the target image; and the text recognition model is used to acquire the text information in the target image, as shown in the figure. Figure 5

[0053] In some embodiments, before the image data in the target image is acquired, a long strip comic image is cut. By analyzing the pixel value of the image, it is determined to cut at a non-first image content, and the long strip comic image is cut to obtain a plurality of short strip comic images, and then the short strip comic images are determined as target images.

[0054] In one example, a strip comic image with a length of 20000 is cut to obtain a plurality of target images with a length less than 1000. The specific length can be set according to actual conditions, and no limitation is made thereto.

[0055] S120, based on a preset rule and the position relationship information of the first text information and the first image, the first text information and the first image are merged to obtain a second image.

[0056] ​​​​​​In some embodiments, the preset rule comprises a rule of determining the merging according to the position relationship between the first text information and the first image. The preset rule is used to determine whether to merge the first text information and the first image to obtain the second image.

[0057] In some embodiments, the merging of the first text information and the first image to obtain the second image based on the preset rule and the position relationship information between the first text information and the first image comprises: when the first image and a plurality of first images coincide or partially coincide with a preset area where the first image extends in a preset direction, merging the first image and the plurality of first images to obtain the second image, as shown in Figure 6 In some embodiments, the merging of the first text information and the first image to obtain the second image based on the preset rule and the position relationship information between the first text information and the first image comprises: when the first text information and the first image overlap, merging the first text information and the first image to obtain the second image, as shown in Figure 2 In some embodiments, the merging of the first text information and the first image to obtain the second image based on the preset rule and the position relationship information between the first text information and the first image comprises: when the first text information and the first image partially overlap, merging the first text information and the first image to obtain the second image, as shown in Figure 3 In some embodiments, the merging of the first text information and the first image to obtain the second image based on the preset rule and the position relationship information between the first text information and the first image comprises: when the first text information and the first image do not overlap, and the first text information partially coincides with the preset area, merging the first text information and the first image to obtain the second image, as shown in Figure 7 In some embodiments, the merging of the first text information and the first image to obtain the second image based on the preset rule and the position relationship information between the first text information and the first image comprises: when the first text information and the first image do not overlap, the first text information does not coincide with the preset area, and the distance between the first image and the first text information is less than a first threshold, merging the first image and the first text information to obtain the second image, as shown in Figure 8 In some embodiments, the merging of the first text information and the first image to obtain the second image based on the preset rule and the position relationship information between the first text information and the first image comprises: when the first text information and the first image do not overlap, the first text information does not coincide with the preset area, and the distance between the first image and the first text information is greater than the first threshold, determining the first text information as the second image, as shown in Figure 9 In some embodiments, the merging of the first text information and the first image to obtain the second image based on the preset rule and the position relationship information between the first text information and the first image comprises: when the first text information and the first image do not overlap, the first text information does not coincide with the preset area, the distance between the first image and the first text information is greater than the first threshold, and the distance between the first image and third text information is greater than a third threshold, merging the third text information and the first text information to obtain the second image, as shown in Figure 10 In some embodiments, the merging of the first text information and the first image to obtain the second image based on the preset rule and the position relationship information between the first text information and the first image comprises: when the first text information and the first image do not overlap, the first text information does not coincide with the preset area, the distance between the first image and the first text information is greater than the first threshold, and the distance between the first image and third text information is less than the third threshold, merging the first text information, the third text information and the first image to obtain the second image, as shown in Figure 11 In some embodiments, the merging of the first text information and the first image to obtain the second image based on the preset rule and the position relationship information between the first text information and the first image comprises: when the first text information and the first image do not overlap, the first text information does not coincide with the preset area, and the distance between the first image and the first text information is less than a first threshold, merging the first image and the first text information to obtain the second image, as shown in

[0058] S130, splitting the target image into a plurality of third images based on the second image and a preset size.

[0059] In some embodiments, the target image is split into a plurality of second images, when a size of a second image is less than or equal to a preset size, the second image is determined as a third image; when the size of the second image is greater than the preset size, the second image is split to obtain a plurality of third images less than or equal to the preset size, as shown in Figure 12

[0060] In an example, the preset size is 2 / 3 of the display area of the user client, and the specific size can be set according to actual needs, which is not limited.

[0061] S140, obtaining image data of the third image, the image data comprising second text information, the second text information comprising text and text density.

[0062] In some embodiments, the image data of the third image is obtained by an AI model. The second text information can further comprise basic information and density of barrage information. The density of barrage information is calculated by the number of barrages in the third image and a barrage coefficient; the basic information is a basic dwell time of the third image, when the third image is pure text, the basic dwell time is 0, and when the third image is non-pure text, the basic dwell time is not 0, which can be selected according to actual conditions, which is not limited.

[0063] S150, determining a speed of automatic playing of the third image based on the second text information of the third image.

[0064] In some embodiments, the speed of automatic playing of the third image is determined based on the second text information of the third image, comprising: determining a text weight according to the second text information; determining the speed of automatic playing of the third image according to the text weight and the text density. The greater the text density, the slower the speed of automatic playing of the third image; the smaller the text density, the faster the speed of automatic playing of the third image. The text density calculation formula is as follows:

[0065] Text density = text weight x text quantity

[0066] In some embodiments, the text weight is determined according to the text in the third image, comprising: all text initial weights are 1, when it is identified that the text in the third image contains a preset keyword, the weight of the text is reduced.

[0067] In some embodiments, when the second text information comprises basic information and density of barrage information, the text information density calculation formula is as follows:

[0068] Text information density = basic information + (text weight x text quantity) + (barrage coefficient x barrage quantity)

[0069] S160, playing the third image based on the speed of automatic playing.

[0070] ​The method for automatically playing a comic provided in the embodiments of the present application can acquire first images and text information of a target image, and merge the first images according to their position relationship. Then, the target image is split according to the merged image, so that a user can read the complete image content. The playing speed is determined by the text information density in the split image, the playing speed can be adjusted for different images, and the user's reading experience is improved.

[0071] Based on the method for automatically playing a comic, the embodiments of the present application further provide a device for automatically playing a comic, and the specific content is as follows.

[0072] Figure 13 is a device structure schematic diagram provided in the embodiments of the present application. As shown in Figure 13 , the device can include a first acquisition module 1310, a merging module 1320, a splitting module 1330, a second acquisition module 1340, a determination module 1350, and a playing module 1360.

[0073] The first acquisition module 1310 is configured to acquire image data in a target image, the image data including first text information, a plurality of first images, and position relationship information of the first text information and the plurality of first images. The position relationship information includes that the first text information overlaps the first images, or the first text information and the plurality of first images partially overlap, or the first text information and the plurality of first images do not overlap.

[0074] The merging module 1320 is configured to merge the first text information and the first images to obtain a second image based on a preset rule and the position relationship information of the first text information and the first images.

[0075] The splitting module 1330 is configured to split the target image to obtain a plurality of third images based on the second image and a preset size.

[0076] The second acquisition module 1340 is configured to acquire image data of the third image, the image data including second text information, and the second text information including text and text density.

[0077] The determination module 1350 is configured to determine a speed of automatic playing of the third image based on the second text information of the third image.

[0078] The playing module 1360 is configured to play the third image based on the speed of automatic playing.

[0079] The device for automatically playing a comic 1300 provided in the embodiments of the present application acquires first images and text information in a target image, merges the first images according to their position relationship, and then splits the target image according to the merged image, so that a user can read the complete image. The speed of automatic playing of the comic is determined by the text density in the image, the playing speed can be adaptively adjusted for different images, and the user can obtain better reading experience.

[0080] In some embodiments, the first text information and the first image are merged to obtain a second image based on preset rules and position relationship information between the first text information and the first image, specifically including: when the first image and a plurality of first images coincide or partially coincide with a preset area in which the first image is extended in a preset direction, the first image and the plurality of first images are merged to obtain the second image; when the first text information overlaps the first image, the first text information and the first image are merged to obtain the second image; when the first text information and the first image partially overlap, the first text information and the first image are merged to obtain the second image; when the first text information does not overlap the first image, and the first text information partially coincides with the preset area, the first text information and the first image are merged to obtain the second image; when the first text information does not overlap the first image, the first text information partially coincides with the preset area, and the distance between the first image and the first text information is less than a first threshold, the first image and the first text information are merged to obtain the second image; when the first text information does not overlap the first image, the first text information does not coincide with the preset area, and the distance between the first image and the first text information is greater than the first threshold, the first text information is determined as the second image; when the first text information does not overlap the first image, the first text information does not coincide with the preset area, the distance between the first image and the first text information is greater than the first threshold, and the distance between the first image and third text information is greater than a third threshold, the third text information and the first text information are merged to obtain the second image; when the first text information does not overlap the first image, the first text information does not coincide with the preset area, the distance between the first image and the first text information is greater than the first threshold, and the distance between the first image and the third text information is less than the third threshold, the first text information, the third text information and the first image are merged to obtain the second image; wherein the third text information is text information with a distance less than a second threshold from the first text information.

[0081] In some embodiments, the speed of automatic playing of the third image is determined based on the second text information of the third image, including: determining a text weight according to the second text information; and determining the speed of automatic playing of the third image according to the text weight and the character density.

[0082] It should be noted that the device of the embodiment can be used as the execution subject of the method of each embodiment, and can implement the corresponding processes in each method to achieve the same technical effects. For brevity, this aspect will not be described in detail.

[0083] Figure 14 A hardware structure schematic diagram of automatic playing of a comic provided by an embodiment of the present application is shown.

[0084] The comic automatic playing device can include a processor 1401 and a memory 1402 storing computer program instructions.

[0085] Specifically, the processor 1401 can include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or can be configured to implement one or more integrated circuits that embody the embodiments of the present application.

[0086] The memory 1402 can include a mass storage for data or instructions. By way of example and not limitation, the memory 1402 can include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive or a combination of two or more of these. In one example, the memory 302 can include removable or non-removable (or fixed) media, or the memory 1402 is a non-volatile solid-state memory. The memory 1402 can be internal or external to the integrated gateway disaster recovery device.

[0087] In one example, the memory 1402 can include read-only memory (ROM), random access memory (RAM), a disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory 1402 includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software that, when executed (e.g., by one or more processors), is operable to perform the operations described with reference to the methods according to an aspect of the present application.

[0088] The processor 1401 implements the methods / step S110 to S160 in the embodiments shown by reading and executing the computer program instructions stored in the memory 1402, and achieves Figure 1 the corresponding technical effects achieved by the methods / steps S110 to S160 in the embodiments shown, which are not described in detail here for brevity. Figure 1

[0089] In one example, the comic automatic playing device can also include a communication interface 1403 and a bus 1410. Wherein, as Figure 14 shown, the processor 1401, the memory 1402, the communication interface 1403 are connected through the bus 1410 and complete the communication between each other.

[0090] The communication interface 1403 is mainly used to realize the communication between each module, device, unit and / or equipment in the embodiments of the present application.

[0091] ​Bus 1410 includes hardware, software, or both, to couple components of the comic automatic playing device to each other and to couple components to other components, such as one or more other devices. While bus 1410 is shown for the sake of clarity as a single bus, bus 1410 can include one or more buses operating together. Bus 1410 can include software configured to implement a communication protocol, such as a Peripheral Component Interconnect (PCI) protocol, a HyperTransport (HT) protocol, a Wireless Bandwidth (WiB) protocol, a Universal Serial Bus (USB) protocol, a Bluetooth® protocol, an Ethernet protocol, a Fiber Channel protocol, or a protocol suitable for communicating data between components of the comic automatic playing device. In some embodiments, bus 1410 can include one or more buses according to various bus standards and protocols. Although bus 1410 is shown as a single bus that interconnects all of the components of the comic automatic playing device, alternative embodiments of the comic automatic playing device can use multiple buses. Bus 1410, or various buses thereof, can be implemented using any suitable type, including a system bus or any other type of interconnection bus or hardware coupled to the components of the comic automatic playing device.

[0092] The comic automatic playing device can perform the comic automatic playing method according to the merged second image and the text information density, thereby realizing the comic automatic playing method according to the comic automatic playing device Figure 1 The comic automatic playing method is described.

[0093] In addition, the comic automatic playing method according to the comic automatic playing device can be implemented by a computer storage medium. The computer storage medium stores computer program instructions. When the computer program instructions are executed by a processor, the comic automatic playing method according to any one of the comic automatic playing device is realized.

[0094] It should be understood that the comic automatic playing device is not limited to the particular configurations and processes described above and shown in the drawings. For the sake of brevity, conventional techniques and methods related to the comic automatic playing device are not described in detail herein. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the comic automatic playing device is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications and additions, or change the order of the steps, after understanding the spirit of the comic automatic playing device.

[0095] The functions noted in the description of the structure block diagrams above can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, and the like. When implemented in software, the elements of the present application are program or code segments that are used to perform the required tasks. The program or code segments can be stored in a machine-readable medium, or transmitted through a data signal carried in a carrier wave over a transmission medium or communication link. A "machine-readable medium" includes any medium that can store or transport information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, and the like. The code segments can be downloaded via computer networks such as the Internet, intranets, and the like.

[0096] It is also important to note that the examples mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the steps mentioned above, that is, the steps can be performed in the order mentioned in the examples, or in an order different from the examples, or several steps can be performed simultaneously.

[0097] The above-described aspects of the present application can be implemented in a computer program product, an apparatus (system), and / or a method of any combination of the above. The computer program product can be a computer storage medium readable by a computer system and encoding a computer program of instructions for executing application programs. The computer storage medium can include, but is not limited to, floppy disks, optical disks, CD-ROMs, magneto-optical disks, ROMs, RAMs, erasable programmable ROMs (EPROMs), electrically erasable programmable ROMs (EEPROMs), magnetic or optical cards, flash memory, and the like.

[0098] The above merely describes a specific implementation of the present application. Those skilled in the art can clearly understand the specific working processes of the system, modules and units described above for the convenience and brevity of description, and can refer to the corresponding processes in the foregoing method embodiments, which will not be described herein again. It should be understood that the protection scope of the present application is not limited to this, and any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered within the protection scope of the present application.

Claims

1. A method for automatically playing a cartoon, characterized by, include: Image data within a target image is acquired. The image data includes first text information, multiple first images, and positional relationship information between the first text information and the multiple first images. The positional relationship information includes whether the first text information overlaps with the first images, or whether the first text information partially overlaps with the multiple first images, or whether the first text information does not overlap with the multiple first images. Based on preset rules and the positional relationship between the first text information and the first image, the first text information and the first image are merged to obtain the second image; The preset rules include: determining the merging rules based on the positional relationship between the first text information and the first image, wherein the preset rules are used to determine whether to merge the first text information and the first image to obtain a second image; Multiple third images are obtained by splitting the target image based on the second image and a preset size; Acquire image data of a third image, wherein the image data includes second text information, and the second text information includes text and character density; The speed at which the third image is automatically played is determined based on the second text information of the third image. The third image is played based on the autoplay speed.

2. The method of claim 1, wherein, The step of merging the first text information and the first image to obtain the second image based on preset rules and the positional relationship information between the first text information and the first image includes: When multiple first images overlap or partially overlap with the preset area where the first image is located, the first image and the multiple first images are merged to obtain a second image, where the preset area is the area covered by the first image after it is extended in a preset direction. When the first text information overlaps with the first image, the first text information and the first image are merged to obtain the second image; When the first text information and the first image partially overlap, the first text information and the first image are merged to obtain the second image; When the first text information does not overlap with the first image, and the first text information partially overlaps with a preset area, the first text information and the first image are merged to obtain the second image; When the first text information does not overlap with the first image, the first text information partially overlaps with a preset area, and the distance between the first image and the first text information is less than a first threshold, the first image and the first text information are merged to obtain the second image. When the first text information does not overlap with the first image, the first text information does not coincide with the preset area, and the distance between the first image and the first text information is greater than the first threshold, the first text information is determined to be the second image. When the first text information does not overlap with the first image, the first text information does not coincide with the preset area, the distance between the first image and the first text information is greater than a first threshold, and the distance between the first image and the third text information is greater than a third threshold, the third text information and the first text information are merged to obtain the second image. When the first text information does not overlap with the first image, the first text information does not coincide with the preset area, the distance between the first image and the first text information is greater than a first threshold, and the distance between the first image and the third text information is less than a third threshold, the first text information, the third text information and the first image are merged to obtain the second image. The third text information is text information whose distance from the first text information is less than the second threshold.

3. The method according to claim 1 or 2, characterized in that, The method of determining the autoplay speed of the third image based on the second text information of the third image includes: The text weight is determined based on the second text information; The autoplay speed of the third image is determined based on the text weight and the text density.

4. An automatic comic playback device, characterized in that, The device includes: The first acquisition module is used to acquire image data within the target image. The image data includes first text information, multiple first images, and positional relationship information between the first text information and the multiple first images. The positional relationship information includes whether the first text information overlaps with the first image, or whether the first text information partially overlaps with the multiple first images, or whether the first text information does not overlap with the multiple first images. A merging module is used to merge the first text information and the first image to obtain a second image based on preset rules and the positional relationship information between the first text information and the first image; the preset rules include: determining the merging rules according to the positional relationship between the first text information and the first image, and the preset rules are used to determine whether to merge the first text information and the first image to obtain a second image; A splitting module is used to split the target image based on a second image and a preset size to obtain multiple third images; The second acquisition module is used to acquire image data of the third image, the image data including second text information, the second text information including text and character density; The determining module is used to determine the automatic playback speed of the third image based on the second text information of the third image; A playback module is used to play the third image based on the autoplay speed.

5. The apparatus according to claim 4, characterized in that, The step of merging the first text information and the first image to obtain the second image based on preset rules and the positional relationship information between the first text information and the first image specifically includes: When multiple first images overlap or partially overlap with the preset area where the first image is located, the first image and the multiple first images are merged to obtain a second image, where the preset area is the area covered by the first image after it is extended in a preset direction. When the first text information overlaps with the first image, the first text information and the first image are merged to obtain the second image; When the first text information and the first image partially overlap, the first text information and the first image are merged to obtain the second image; When the first text information does not overlap with the first image, and the first text information partially overlaps with a preset area, the first text information and the first image are merged to obtain the second image; When the first text information does not overlap with the first image, the first text information does not coincide with the preset area, and the distance between the first image and the first text information is less than the first threshold, the first image and the first text information are merged to obtain the second image; When the first text information does not overlap with the first image, the first text information does not coincide with the preset area, and the distance between the first image and the first text information is greater than the first threshold, the first text information is determined to be the second image. When the first text information does not overlap with the first image, the first text information does not coincide with the preset area, the distance between the first image and the first text information is greater than a first threshold, and the distance between the first image and the third text information is greater than a third threshold, the third text information and the first text information are merged to obtain the second image. When the first text information does not overlap with the first image, the first text information does not coincide with the preset area, the distance between the first image and the first text information is greater than a first threshold, and the distance between the first image and the third text information is less than a third threshold, the first text information, the third text information and the first image are merged to obtain the second image. The third text information is text information whose distance from the first text information is less than the second threshold.

6. The apparatus according to claim 4 or 5, characterized in that, The method of determining the autoplay speed of the third image based on the second text information of the third image includes: The text weight is determined based on the second text information; The autoplay speed of the third image is determined based on the text weight and the text density.

7. An automatic comic playback device, characterized in that, The device includes: a processor, and a memory storing computer program instructions; The processor reads and executes the computer program instructions to implement the automatic comic playback method as described in any one of claims 1-3.

8. A computer storage medium, characterized in that, The computer storage medium stores computer program instructions, which, when executed by a processor, implement the automatic comic playback method as described in any one of claims 1-3.

Citation Information

Patent Citations

  • Cartoon synthesis method, device and apparatus and computer readable storage medium

    CN108320319A

  • Text image processing method and device, electronic scanning equipment and storage medium

    CN111832551A