Text content determination method and apparatus, and readable storage medium

CN118447492BActive Publication Date: 2026-09-25CHINA UNITED NETWORK COMM GRP CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310118990.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-06
Publication Date
2026-09-25
Estimated Expiration
2043-02-06

AI Technical Summary

Technical Problem

但在某些固定的场景,算法需要通过视频图像的前后多帧连续数据来综合计算,比如通过一段时间内视频图像的变换来计算结果,这种情况下视频图像中的数字在动态变化中会存在乱码的情况,很容易将乱码部分数字漏识别或者识别错误,导致数字识别的准确率较低

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118447492B_ABST
    Figure CN118447492B_ABST
Patent Text Reader

Abstract

The application discloses a text content determination method and device and a readable storage medium, relates to the technical field of data processing, and is used for improving the recognition accuracy of text content. The method comprises the following steps: acquiring to-be-recognized video data; the to-be-recognized video data comprises a plurality of image frames; a perspective image of a target region in each image frame is determined; the perspective image of the target region in a first image frame is detected, the recognition result of one or more target type texts in the first image frame and the width of the target region in the first image frame are determined, and the recognition result of the one or more target type texts in each image frame is obtained; based on the number of the one or more target type texts and the width of the target region in each image frame, an effective image frame is determined from the plurality of image frames; the effective image frame is an image frame not comprising random codes; and based on the recognition result of the one or more target type texts in the plurality of effective image frames, the text content in the to-be-recognized video data is determined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus and readable storage medium for determining text content. Background Technology

[0002] With the increasing demand for intelligent technology, more and more video images need to be digitally recognized, making it crucial to improve the accuracy of digital recognition.

[0003] Some existing digit recognition methods have achieved high accuracy for static digit recognition. However, in certain fixed scenarios, algorithms need to comprehensively calculate based on continuous data from multiple frames of a video image, such as calculating the result based on the changes in the video image over a period of time. In this case, the digits in the video image may contain garbled characters during dynamic changes, making it easy to miss or misidentify the garbled parts of the digits, resulting in low accuracy in digit recognition. Summary of the Invention

[0004] This application provides a method, apparatus, and readable storage medium for determining text content, used to improve overlapping coverage.

[0005] To achieve the above objectives, this application adopts the following technical solution:

[0006] In a first aspect, a method for determining text content is provided, comprising: acquiring video data to be recognized; the video data to be recognized includes multiple image frames; determining the perspective image of a target region in each image frame; the target region is the text display area in each image frame; detecting the perspective image of the target region in the first image frame, determining the recognition result of one or more target type texts in the first image frame, and the width of the target region in the first image frame, to obtain the recognition result of one or more target type texts in each image frame; the first image frame is any one of the multiple image frames; determining a valid image frame from the multiple image frames based on the number of one or more target type texts and the width of the target region in each image frame; the valid image frame is an image frame that does not contain garbled characters; and determining the text content in the video data to be recognized based on the recognition result of one or more target type texts in the multiple valid image frames.

[0007] Optionally, determining the perspective image of the target region in each image frame includes: determining the vertex coordinates of the target region in the first image frame and performing a projection transformation on the vertex coordinates; processing the projection-transformed vertex coordinates using a first transformation function to determine the projection transformation matrix of the vertex coordinates; and processing the projection transformation matrix using a second transformation function to determine the perspective image of the target region in the first image frame, thereby obtaining the perspective image of the target region in each image frame.

[0008] Optionally, determining the width of the target region in the first image frame includes: determining the maximum and minimum values ​​of the horizontal coordinates of the vertex coordinates of the target region in the first image frame; and determining the width of the target region in the first image frame based on the difference between the maximum and minimum values ​​of the horizontal coordinates.

[0009] Optionally, based on the number of one or more target type texts and the width of the target region in each image frame, a valid image frame is determined from multiple image frames, including: determining the target width range of the target region according to the number of one or more target type texts and a first mapping relationship; the first mapping relationship includes different numbers and corresponding target width ranges; if the width of the target region is within the target width range, the first image frame is determined as a valid image frame, thus obtaining multiple valid image frames.

[0010] Based on the technical solution provided in this application, after acquiring the video data to be recognized, the determining device can determine the perspective image of the target region in each image frame based on multiple image frames in the video data to be recognized. Due to the characteristics of the perspective image, the determining device can more accurately determine the recognition result of one or more target type texts in the first image frame and the width of the target region in the first image frame. Furthermore, based on the number of one or more target type texts and the width of the target region in each image frame, the determining device determines valid image frames from multiple image frames; valid image frames are those that do not contain garbled characters. In this way, the determining device only recognizes image frames that do not contain garbled characters, avoiding misidentification of garbled characters as other text. Therefore, when determining the text content in the video data to be recognized based on the recognition results of one or more target type texts in multiple valid image frames, the recognition accuracy of multiple consecutive image frames can be further improved.

[0011] Secondly, a device for determining text content is provided. The device includes: an acquisition unit and a determination unit. The acquisition unit is used to acquire video data to be recognized; the video data to be recognized includes multiple image frames. The determination unit is used to determine the perspective image of a target region in each image frame; the target region is the text display area in each image frame. The determination unit is further used to detect the perspective image of the target region in the first image frame, determine the recognition result of one or more target type texts in the first image frame, and the width of the target region in the first image frame, to obtain the recognition result of one or more target type texts in each image frame; the first image frame is any one of the multiple image frames. The determination unit is further used to determine a valid image frame from the multiple image frames based on the number of one or more target type texts and the width of the target region in each image frame; the valid image frame is an image frame that does not contain garbled characters. The determination unit is further used to determine the text content in the video data to be recognized based on the recognition result of one or more target type texts in the multiple valid image frames.

[0012] Optionally, the MR data also includes a time advance (TA) value. The determination unit is specifically used to: determine the vertex coordinates of the target region in the first image frame and perform a projection transformation on the vertex coordinates; process the projection-transformed vertex coordinates using a first transformation function to determine the projection transformation matrix of the vertex coordinates; process the projection transformation matrix using a second transformation function to determine the perspective image of the target region in the first image frame, thereby obtaining the perspective image of the target region in each image frame.

[0013] Optionally, the determining unit is further configured to: determine the maximum and minimum values ​​of the horizontal coordinates of the vertex coordinates of the target region in the first image frame; and determine the width of the target region in the first image frame based on the difference between the maximum and minimum values ​​of the horizontal coordinates.

[0014] Optionally, the determining unit is further configured to: determine the target width range of the target region based on the number of one or more target type texts and a first mapping relationship; the first mapping relationship includes different numbers and corresponding target width ranges; and when the width of the target region is within the target width range, determine the first image frame as a valid image frame to obtain multiple valid image frames.

[0015] Thirdly, a text content determination device is provided. This text content determination device can realize the functions performed by the text content determination device in the above-mentioned aspects or possible designs. The functions can be implemented by hardware. For example, in one possible design, the text content determination device may include a processor and a communication interface. The processor can be used to support the text content determination device in realizing the functions involved in the first aspect or any possible design of the first aspect.

[0016] In another possible design, the text content determination device may further include a memory for storing necessary computer execution instructions and data for the text content determination device. When the text content determination device is running, the processor executes the computer execution instructions stored in the memory to cause the text content determination device to perform the first aspect or any of the possible text content determination methods described above.

[0017] Fourthly, a computer-readable storage medium is provided, which may be a readable non-volatile storage medium storing computer instructions or programs that, when executed on a computer, enable the computer to perform the method for determining the text content described in the first aspect or any of the possible methods described above.

[0018] Fifthly, a computer program product containing instructions is provided, which, when run on a computer, enables the computer to perform the method for determining text content according to the first aspect or any possible design of the above aspects.

[0019] A sixth aspect provides an electronic device comprising one or more processors and one or more memories. The one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer program code, including computer instructions, which, when executed by the one or more processors, cause the electronic device to perform a method for determining text content as described in the first aspect or any possible design of the first aspect.

[0020] In a seventh aspect, a chip system is provided, comprising a processor and a communication interface, which can be used to implement the functions performed by the means for determining text content in the first aspect or any possible design of the first aspect. In one possible design, the chip system further includes a memory for storing program instructions and / or data. The chip system may be composed of chips or may include chips and other discrete devices, without limitation. Attached Figure Description

[0021] Figure 1 A schematic diagram of a text content determination system provided in an embodiment of this application;

[0022] Figure 2 A schematic diagram of the structure of a text content determination device provided in an embodiment of this application;

[0023] Figure 3 A flowchart illustrating a method for determining text content provided in an embodiment of this application;

[0024] Figure 4 A flowchart illustrating yet another method for determining text content provided in an embodiment of this application;

[0025] Figure 5 A flowchart illustrating yet another method for determining text content provided in an embodiment of this application;

[0026] Figure 6 A flowchart illustrating yet another method for determining text content provided in an embodiment of this application;

[0027] Figure 7 This is a schematic diagram of the structure of another text content determination device provided in an embodiment of this application. Detailed Implementation

[0028] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0029] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.

[0030] It should also be understood that the term "comprising" indicates the presence of the described feature, whole, step, operation, element and / or component, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements and / or components.

[0031] With the increasing demand for intelligent technology, more and more video images need to be digitally recognized, making it crucial to improve the accuracy of digital recognition.

[0032] Existing digit recognition methods have achieved high accuracy for static digit recognition. However, in certain fixed scenarios, algorithms need to comprehensively calculate using continuous data from multiple frames of a video image, such as calculating results based on the changes in the video image over a period of time. In this case, the digits in the video image may contain garbled characters during dynamic changes, easily leading to missed or incorrect recognition of the garbled parts, resulting in low digit recognition accuracy. Therefore, this application provides a method for determining text content, including:

[0033] Acquire video data to be recognized; the video data to be recognized includes multiple image frames; determine the perspective image of the target region in each image frame; the target region is the text display area in each image frame; detect the perspective image of the target region in the first image frame, determine the recognition result of one or more target type texts in the first image frame, and the width of the target region in the first image frame, to obtain the recognition result of one or more target type texts in each image frame; the first image frame is any one of the multiple image frames; based on the number of one or more target type texts and the width of the target region in each image frame, determine the valid image frames from the multiple image frames; the valid image frames are image frames that do not contain garbled characters; based on the recognition result of one or more target type texts in the multiple valid image frames, determine the text content in the video data to be recognized.

[0034] The methods provided in the embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0035] It should be noted that the network system described in the embodiments of this application is for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and does not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of network systems and the emergence of other network systems, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0036] Figure 1 The diagram shown is a schematic representation of a text content determination system provided in an embodiment of this application. Figure 1 As shown, the system for determining text content may include a text content determining device 11 (hereinafter referred to as the determining device) and a data acquisition device 12. The determining device 11 and the data acquisition device 12 are connected. The determining device 11 and the data acquisition device 12 can be connected wirelessly.

[0037] The determining device 11 involved in the embodiments of this application may also be referred to as a computer, server, etc. In the embodiments of this application, the specific technology and equipment form used for the determining device 11 are not limited.

[0038] In the embodiments of this application, the acquisition device 12 may be a device with video shooting function such as a camera. The embodiments of this application do not limit the specific technology, quantity and form of the acquisition device 12.

[0039] The acquisition device 12 is used to capture images of the target scene, obtain video data of the target scene to be identified, and send the video data of the target scene to the determining device 11. The determining device 11 is used to receive the video data of the target scene to be identified sent by the acquisition device 12, and determine the text content in the video data to be identified based on the video data of the target scene.

[0040] In different application scenarios, the data acquisition device 12 and the determination device 11 can be independent devices or integrated into the same device. This embodiment of the invention does not impose specific limitations on this.

[0041] It should be noted that, Figure 1 This is just an example framework diagram. Figure 1 The names of the various devices included are unrestricted, and except for Figure 1 In addition to the functional nodes shown, other nodes may also be included, but this application embodiment does not limit this.

[0042] It should be noted that, Figure 1 This is just an example framework diagram. Figure 1The names of the modules included are unrestricted, and except for Figure 1 In addition to the functional modules shown, other modules may also be included, but this application embodiment does not limit this.

[0043] In practical implementation, Figure 1 Each device in the process can be adopted Figure 2 The shown composition structure, or including Figure 2 The components shown. Figure 2 This is a schematic diagram illustrating the composition of a determining device 200 provided in an embodiment of this application. The determining device 200 can be a server, or it can be a chip or system-on-a-chip within the server. Figure 2 As shown, the determining device 200 includes a processor 201, a communication interface 202, and a communication line 203.

[0044] Furthermore, the determining device 200 may also include a memory 204. The processor 201, memory 204, and communication interface 202 can be connected via a communication line 203.

[0045] The processor 201 can be a CPU, a general-purpose processor, a network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. The processor 201 can also be other devices with processing capabilities, such as circuits, devices, or software modules, without limitation.

[0046] Communication interface 202 is used to communicate with other devices or other communication networks. Communication interface 202 can be a module, circuit, communication interface, or any device capable of enabling communication.

[0047] Communication line 203 is used to transmit information between the components included in determining device 200.

[0048] Memory 204 is used to store instructions. These instructions can be computer programs.

[0049] The memory 204 can be a read-only memory (ROM) or other type of static storage device that can store static information and / or instructions; it can also be a random access memory (RAM) or other type of dynamic storage device that can store information and / or instructions; it can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, etc., without limitation.

[0050] It should be noted that the memory 204 can exist independently of the processor 201 or can be integrated with the processor 201. The memory 204 can be used to store instructions, program code, or some data, etc. The memory 204 can be located inside or outside the determining device 200, without limitation. The processor 201 is used to execute the instructions stored in the memory 204 to implement the text content determination method provided in the following embodiments of this application.

[0051] In one example, processor 201 may include one or more CPUs, for example, Figure 2 CPU0 and CPU1 in the CPU.

[0052] As an optional implementation, the determining device 200 includes multiple processors, for example, besides Figure 2 In addition to processor 201, it may also include processor 205.

[0053] It should be pointed out that, Figure 2 The composition shown does not constitute a basis for the interpretation of this invention. Figure 1 The limitations of each device in the process, except Figure 2 In addition to the components shown, Figure 1 The various devices in the can include ratio Figure 2 More or fewer components, or combinations of certain components, or different arrangements of components.

[0054] In this embodiment of the application, the chip system may be composed of chips or may include chips and other discrete devices.

[0055] Furthermore, the actions, terms, etc., involved in the various embodiments of this application can be referenced interchangeably without limitation. The message names or parameter names in the messages exchanged between the various devices in the embodiments of this application are merely examples, and other names may be used in specific implementations without limitation.

[0056] To facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.

[0057] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0058] In this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0059] The following is combined with Figure 1 The system for determining text content shown describes the method for determining text content provided in the embodiments of this application.

[0060] Figure 3 This application provides a method for determining text content, which can be applied to a server or a determining device. Figure 1 The determining device 11 can also be a component within the determining device 11, such as a chip. This application embodiment uses an application to the determining device 11 as an example for illustration. Figure 3 As shown, the method includes the following steps S301-S304:

[0061] S301, The device acquires the video data to be identified.

[0062] The video data to be identified includes multiple image frames; each image frame may include a display screen. For example, the display screen can be a digital tube display, an LED display, etc. The display screen may include multiple types of text, such as numbers, text, and random characters.

[0063] As one possible implementation, the determining device can be communicatively connected to the acquiring device. The determining device can send a first request message to the acquiring device, requesting video data of the target scene. After receiving the first request message from the determining device, the acquiring device can send a first reply message to the determining device, which includes the video data of the target scene. Accordingly, the determining device can obtain the video data of the target scene by receiving the first reply message from the acquiring device.

[0064] S302, The determining device determines the perspective image of the target region in each image frame.

[0065] The target region is the text display area in each image frame. For example, it can be a display screen area, or it can be the smallest bounding rectangle of the text display area. The size of the perspective image is the same as the size of the source image.

[0066] As one possible implementation, the determining device can determine the projection transformation matrix of the vertex coordinates of the target region in each image frame, and process the projection transformation matrix using a perspective transformation function to determine the perspective image of the target region in each image frame.

[0067] It should be noted that the specific details of determining the perspective image of the target region in each image frame in this possible implementation will be explained in a later section, and will not be repeated here.

[0068] S303, The determining device detects the perspective image of the target region in the first image frame, determines the recognition result of one or more target type texts in the first image frame, and the width of the target region in the first image frame, and obtains the recognition result of one or more target type texts in each image frame.

[0069] The image frame can be any one of multiple image frames. The target type text can be numbers, text, etc., without restriction.

[0070] As one possible implementation, the determining device can employ an Optical Character Recognition (OCR) algorithm trained with a neural network to detect the perspective image of the target region in the first image frame, obtaining the coordinates of one or more rectangular boxes, and the recognition result of the target type text within each rectangular box. Furthermore, the determining device can determine the width of the target region in the first image frame based on the coordinates of the rectangular boxes in the first image frame.

[0071] It should be noted that, in order to avoid the algorithm misidentifying garbled characters as other text, the target area (also known as the detection box) includes the garbled characters during data annotation, but the recognized content only labels the normally displayed numbers and does not label the garbled characters.

[0072] S304. The determining device determines the valid image frame from multiple image frames based on the number of one or more target type texts and the width of the target region in each image frame.

[0073] Among them, a valid image frame is an image frame that does not contain garbled characters.

[0074] As one possible implementation, the determining device can pre-store the width range of the target region corresponding to the number of different target type texts. The determining device can determine the first image frame as a valid image frame if the number of one or more target type texts in the first image frame corresponds to the width of the target region, and then determine valid image frames from multiple image frames.

[0075] It should be noted that the specific details of determining the valid image frame from multiple image frames based on the number of one or more target type texts and the width of the target region in each image frame will be explained in a later section, and will not be repeated here.

[0076] S305. The determining device determines the text content in the video data to be recognized based on the recognition results of one or more target type texts in multiple valid image frames.

[0077] As one possible implementation, the determining device can output the recognition results of multiple valid image frames in sequence according to the playback time of the video data, thereby obtaining the text content in the video data to be recognized.

[0078] As one possible implementation, the determining device can output the recognition results of multiple valid image frames in sequence according to the playback time of the video data, and merge valid images with the same and consecutive recognition results to obtain the text content in the video data to be recognized.

[0079] For example, the recognition results of multiple valid image frames can be 1, 12, 12, 18, 16. Since the second and third valid image frames are consecutive and have the same recognition result, the text content in the merged video data to be recognized can be 1, 12, 18, 16.

[0080] Based on the technical solution provided in this application, after acquiring the video data to be recognized, the determining device can determine the perspective image of the target region in each image frame based on multiple image frames in the video data to be recognized. Due to the characteristics of the perspective image, the determining device can more accurately determine the recognition result of one or more target type texts in the first image frame and the width of the target region in the first image frame. Furthermore, based on the number of one or more target type texts and the width of the target region in each image frame, the determining device determines valid image frames from multiple image frames; valid image frames are those that do not contain garbled characters. In this way, the determining device only recognizes image frames that do not contain garbled characters, avoiding misidentification of garbled characters as other text. Therefore, when determining the text content in the video data to be recognized based on the recognition results of one or more target type texts in multiple valid image frames, the recognition accuracy of multiple consecutive image frames can be further improved.

[0081] One possible implementation, such as Figure 4 As shown, in order to determine the perspective image of the target region in each image frame, S302 in the determination method of this application may further include the following S401-S403.

[0082] S401. The determining device determines the vertex coordinates of the target region in the first image frame and performs a projection transformation on the vertex coordinates.

[0083] The vertex coordinates of the target region can include the coordinates of the top-left corner, bottom-left corner, top-right corner, and bottom-right corner of the target region. For example, the vertex coordinates of the target region in the first image frame can be (x1, y1), (x2, y2), (x3, y3), and (x4, y4), respectively.

[0084] As one possible implementation, the determining device can input each image frame into the display device detection model to obtain the minimum bounding rectangle of the display device in each image frame, as well as the vertex pixel coordinates of the minimum bounding rectangle. This allows the determination of the vertex coordinates of the target region in the first image frame.

[0085] As another possible implementation, the determining device can input each image frame into the text detection model to obtain the minimum bounding rectangle of the text content in each image frame, as well as the vertex pixel coordinates of the minimum bounding rectangle. This allows the determination of the vertex coordinates of the target region in the first image frame.

[0086] Furthermore, the determining device can determine the vertex coordinates after projection transformation according to the following formulas one and two. For example, the vertex coordinates after projection transformation are (0, 0), (x... p ,0),(x p ,y p (0,y) p In the case of ), Formula 1 and Formula 2 can be:

[0087]

[0088]

[0089] S402. The determining device uses the first transformation function to process the vertex coordinates after projection transformation and determines the projection transformation matrix of the vertex coordinates.

[0090] The first transformation function can be set as needed. For example, it can be the getPerspectiveTransform function from OpenCV, a cross-platform computer vision and machine learning software library.

[0091] As one possible implementation, the determining device can substitute the vertex coordinates after projection transformation into the first transformation function to solve for the projection transformation matrix of the vertex coordinates.

[0092] For example, the first transformation function can be M = getPerspectiveTransform(points1, points2).

[0093] Where M represents the projection transformation matrix of the vertex coordinates. points1 represents the vertex coordinates of the target region in the source image. points2 represents the vertex coordinates after the projection transformation.

[0094] S403. The determining device uses the second transformation function to process the projection transformation matrix, determines the perspective image of the target region in the first image frame, and obtains the perspective image of the target region in each image frame.

[0095] The second transformation function can be set as needed. For example, it can be the warpPerspective function in OpenCV.

[0096] As one possible implementation, the determining device can substitute the projection transformation matrix into the second transformation function for solving, determine the perspective image of the target region in the first image frame, and obtain the perspective image of the target region in each image frame.

[0097] For example, the second transformation function can be Output = warpPerspective(img_box, M, (xp, yp)).

[0098] Where Output represents the perspective image of the target region. im g_bo ox represents the image of the target region in the first image frame.

[0099] One possible implementation, such as Figure 5 As shown, in order to determine the width of the target region in the first image frame, the determination method of this application may further include the following steps S501-S502.

[0100] S501, The determining device determines the maximum and minimum values ​​of the horizontal coordinates of the vertex coordinates of the target region in the first image frame.

[0101] As one possible implementation, the determining device includes a comparator internally, used to compare the magnitudes of different coordinates. The determining device can then use the comparator to determine the maximum and minimum x-coordinate values ​​of the vertex coordinates of the target region in the first image frame.

[0102] In one example, the vertex coordinates of the target region in the first image frame can be (1, 1), (60, 1), (1, 80), and (60, 80). Then the determining device can determine that the maximum value of the x-coordinate of the vertex coordinates of the target region in the first image frame is 60, and the minimum value of the x-coordinate of the vertex coordinates of the target region in the first image frame is 1.

[0103] S502, The determining device determines the width of the target region in the first image frame based on the difference between the maximum value and the minimum value of the horizontal coordinate.

[0104] In one example, if the maximum value of the x-coordinate in the vertex coordinates of the target region in the first image frame is 60 and the minimum value of the x-coordinate in the vertex coordinates of the target region in the first image frame is 1, the determining device can determine that the difference between the maximum value and the minimum value of the x-coordinate is 60-1=59, and determine 59 as the width of the target region in the first image frame.

[0105] One possible implementation, such as Figure 6 As shown, in order to determine the valid image frame from multiple image frames based on the number of one or more target type texts and the width of the target region in each image frame, S304 of the determination method of this application may further include the following S601-S602.

[0106] S601, The determining device determines the target width range of the target region based on the number of one or more target type texts and a first mapping relationship.

[0107] The first mapping relationship includes different quantities and corresponding target width ranges. For example, the first mapping relationship can be shown in Table 1 below.

[0108] Table 1 First Mapping Relationship Table

[0109] 1 18 pixels - 22 pixels 2 38 pixels - 42 pixels 3 58 pixels - 62 pixels 4 78 pixels - 82 pixels

[0110] It should be noted that the data in Table 1 is merely exemplary. In this embodiment, the number of one or more target type texts and the width of the target area can also be other data, without limitation.

[0111] For example, when the number of one or more target type texts is 2, the determining device can determine the target width range of the target region as 38 pixels to 42 pixels based on the first mapping relationship.

[0112] S602. When the width of the target area is within the target width range, the device determines the first image frame as a valid image frame and obtains multiple valid image frames.

[0113] For example, if the target width of the target area is between 38 and 42 pixels, and the target area is 400 pixels wide, the determining device can determine that the first image frame is a valid image frame.

[0114] The number of one or more target type texts and the width of the target region satisfy the first mapping relationship.

[0115] The various solutions in the above embodiments of this application can be combined without contradiction.

[0116] This application embodiment can divide the determining device into functional modules or functional units according to the above method examples. For example, each function can be divided into a separate functional module or functional unit, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or in software functional modules or functional units. The module or unit division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.

[0117] When dividing each function into modules according to its corresponding function. Figure 7 A schematic diagram of a determining device is shown. This determining device can be a server or a chip applied in a server. This determining device can be used to perform the functions related to the server involved in the above embodiments. Figure 7The determining device shown may include: an acquisition unit 701 and a determining unit 702; the acquisition unit 701 is used to acquire video data to be identified; the video data to be identified includes multiple image frames; the determining unit 702 is used to determine the perspective image of a target region in each image frame; the target region is the text display area in each image frame; the determining unit 702 is further used to detect the perspective image of the target region in the first image frame, determine the recognition result of one or more target type texts in the first image frame, and the width of the target region in the first image frame, to obtain the recognition result of one or more target type texts in each image frame; the first image frame is any one of the multiple image frames; the determining unit 702 is further used to determine a valid image frame from the multiple image frames based on the number of one or more target type texts and the width of the target region in each image frame; the valid image frame is an image frame that does not contain garbled characters; the determining unit 702 is further used to determine the text content in the video data to be identified based on the recognition result of one or more target type texts in the multiple valid image frames.

[0118] In one possible design, the MR data also includes a time advance (TA) value. The determination unit 702 is specifically used to: determine the vertex coordinates of the target region in the first image frame and perform a projection transformation on the vertex coordinates; process the projection-transformed vertex coordinates using a first transformation function to determine the projection transformation matrix of the vertex coordinates; process the projection transformation matrix using a second transformation function to determine the perspective image of the target region in the first image frame, thereby obtaining the perspective image of the target region in each image frame.

[0119] In one possible design, the determining unit 702 is further configured to: determine the maximum and minimum values ​​of the horizontal coordinates of the vertex coordinates of the target region in the first image frame; and determine the width of the target region in the first image frame based on the difference between the maximum and minimum values ​​of the horizontal coordinates.

[0120] In one possible design, the determining unit 702 is further configured to: determine the target width range of the target region based on the number of one or more target type texts and a first mapping relationship; the first mapping relationship includes different numbers and corresponding target width ranges; and when the width of the target region is within the target width range, determine the first image frame as a valid image frame, thereby obtaining multiple valid image frames.

[0121] This application also provides a computer-readable storage medium. All or part of the processes in the above method embodiments can be implemented by a computer program instructing related hardware. This program can be stored in the computer-readable storage medium, and when executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be an internal storage unit of the determining device (including a data sending end and / or a data receiving end) in any of the foregoing embodiments, such as a hard disk or memory of the determining device. The computer-readable storage medium can also be an external storage device of the terminal device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the terminal device. Further, the computer-readable storage medium can include both the internal storage unit of the determining device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the determining device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0122] It should be noted that the terms "first" and "second," etc., in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0123] It should be understood that in this application, "at least one (item)" means one or more, "more than one" means two or more, "at least two (items)" means two or three or more, and "and / or" is used to describe the relationship between related objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0124] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0125] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0126] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0127] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0128] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0129] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for determining text content, characterized in that, The method includes: Acquire video data to be identified; the video data to be identified includes multiple image frames; Determine the perspective image of the target region in each image frame; the target region is the text display area in each image frame; The perspective image of the target region in the first image frame is detected to determine the recognition result of one or more target type texts in the first image frame and the width of the target region in the first image frame, so as to obtain the recognition result of one or more target type texts in each image frame; the first image frame is any one of the plurality of image frames, and the one or more target type texts in the first image frame do not include garbled texts; Determining valid image frames from a plurality of image frames based on the number of one or more target type texts and the width of the target region in each image frame includes: determining the target width range of the target region according to the number of one or more target type texts and a first mapping relationship; the first mapping relationship includes different numbers and corresponding target width ranges; if the width of the target region is within the target width range, determining the first image frame as a valid image frame, thereby obtaining a plurality of valid image frames; the valid image frames are image frames that do not contain garbled characters. Based on the recognition results of one or more target type texts in multiple valid image frames, the text content in the video data to be recognized is determined.

2. The method according to claim 1, characterized in that, The step of determining the perspective image of the target region in each image frame includes: Determine the vertex coordinates of the target region in the first image frame, and perform a projection transformation on the vertex coordinates; The projection transformation matrix of the vertex coordinates is determined by processing the projection transformation matrix using the first transformation function. The projection transformation matrix is ​​processed using the second transformation function to determine the perspective image of the target region in the first image frame, thereby obtaining the perspective image of the target region in each image frame.

3. The method according to claim 1, characterized in that, Determining the width of the target region in the first image frame includes: Determine the maximum and minimum values ​​of the x-coordinates of the vertex coordinates of the target region in the first image frame; The width of the target region in the first image frame is determined based on the difference between the maximum and minimum values ​​of the horizontal coordinate.

4. A device for determining text content, characterized in that, The device includes: an acquisition unit and a determination unit; The acquisition unit is used to acquire video data to be identified; the video data to be identified includes multiple image frames; The determining unit is used to determine the perspective image of the target region in each image frame; the target region is the text display area in each image frame. The determining unit is further configured to detect the perspective image of the target region in the first image frame, determine the recognition result of one or more target type texts in the first image frame, and the width of the target region in the first image frame, and obtain the recognition result of one or more target type texts in each image frame; the first image frame is any one of the plurality of image frames, and the one or more target type texts in the first image frame do not include garbled texts; The determining unit is further configured to determine a valid image frame from the plurality of image frames based on the number of one or more target type texts and the width of the target region in each image frame; the valid image frame is an image frame that does not contain garbled characters; The determining unit is further configured to determine the text content in the video data to be identified based on the recognition results of one or more target type texts in the plurality of valid image frames; The determining unit is further configured to: The target width range of the target region is determined based on the quantity of the one or more target type texts and a first mapping relationship; the first mapping relationship includes different quantities and corresponding target width ranges. If the width of the target region is within the target width range, the first image frame is determined to be a valid image frame, and a plurality of valid image frames are obtained.

5. The apparatus according to claim 4, characterized in that, The MR data also includes the timing advance (TA) value. The determining unit is specifically used for: Determine the vertex coordinates of the target region in the first image frame, and perform a projection transformation on the vertex coordinates; The projection transformation matrix of the vertex coordinates is determined by processing the projection transformation matrix using the first transformation function. The projection transformation matrix is ​​processed using the second transformation function to determine the perspective image of the target region in the first image frame, thereby obtaining the perspective image of the target region in each image frame.

6. The apparatus according to claim 4, characterized in that, The determining unit is further configured to: Determine the maximum and minimum values ​​of the x-coordinates of the vertex coordinates of the target region in the first image frame; The width of the target region in the first image frame is determined based on the difference between the maximum and minimum values ​​of the horizontal coordinate.

7. The apparatus according to any one of claims 4-6, characterized in that, The determining unit is further configured to: The target width range of the target region is determined based on the quantity of the one or more target type texts and a first mapping relationship; the first mapping relationship includes different quantities and corresponding target width ranges. If the width of the target region is within the target width range, the first image frame is determined to be a valid image frame, and a plurality of valid image frames are obtained.

8. A computer-readable storage medium, characterized in that, The readable storage medium stores instructions that, when executed, implement the method as described in any one of claims 1-3.

9. A device for determining text content, characterized in that, include: The processor, memory, and communication interface; wherein the communication interface is used for communication between the text content determination device and other devices or networks; The memory is used to store one or more programs, the one or more programs including computer-executable instructions, which, when the text content determination device is running, are executed by the processor to execute the computer-executable instructions stored in the memory to cause the text content determination device to perform the method of any one of claims 1-3.

Citation Information

Patent Citations

  • Video text conversion method, mobile terminal and computer readable storage medium

    CN111832529A

  • Video file word segmentation method and device and electronic equipment

    CN113901816A