Information processing method, information transmission method, electronic device, storage medium and computer program product
By requesting and redrawing text description information from the server via the client, the problem of insufficient text clarity in weak network environments is solved, and the effect of improving text clarity with low resource consumption is achieved.
Patent Information
- Application Number
- CN202511497235.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2026-02-10
AI Technical Summary
In weak network environments, the clarity of text in video footage is insufficient to meet users' viewing needs. Existing technologies such as high-definition video streaming consume large amounts of resources, multi-view video streaming is costly, and image super-resolution processing is ineffective.
During video playback, the client requests text description information from the server, processes the image based on OCR technology and redraws the text, and the server stores and transmits the text description information to support the redrawing.
It improves the clarity of text in video images under weak network conditions, reduces transmission resource consumption, and meets practical application needs.
Smart Images

Figure CN121509732A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video processing, and more particularly to an information processing method, an information transmission method, an electronic device, a storage medium, and a computer program product. Background Technology
[0002] Currently, the following two methods are commonly used to solve the problem of insufficient image clarity in videos (such as live videos, on-demand videos, etc.): the source end (which can also be understood as the video source end) provides higher definition video sources for users to choose from; the source end provides video sources with multiple perspectives to improve the recognizability of local content so that users can choose the viewing perspective according to the content they are interested in.
[0003] However, in practical applications, transmitting high-definition video sources usually requires a large amount of bandwidth. At the same time, in order to provide video sources with multiple perspectives (which can also be understood as multiple streams), the acquisition and distribution costs need to be increased significantly, making it difficult to meet the needs of practical applications. Summary of the Invention
[0004] To address the related technical problems, embodiments of this application provide an information processing method, an information transmission method, an electronic device, a storage medium, and a computer program product.
[0005] The technical solution of this application embodiment is implemented as follows: An information processing method, applied to a client, includes: During video playback, a first message is sent to the server, the first message being used to request information related to the text of the first frame of the video; Receive second information sent by the server, the second information being used to describe the text of the first screen; Using the second information, redraw the text in the first image.
[0006] In the above scheme, sending the first information to the server includes: The third piece of information is determined, which indicates whether the text clarity of the first image meets the viewing requirements; If the third information indicates that the text clarity of the first image does not meet the viewing requirements, the first information is sent to the server.
[0007] In the above scheme, determining the third information includes: The first image is processed based on Optical Character Recognition (OCR) technology to obtain fourth information, which represents the recognition confidence of the text in the first image. The third information is determined using the fourth information and the preset threshold.
[0008] This application also provides an information transmission method applied to a server, including: Receive first information sent by the client during video playback, the first information being used to request relevant information about the text of the first frame of the video; Send a second message to the client. The second message is used to describe the text of the first screen. The second message is used to allow the client to redraw the text of the first screen.
[0009] The method in the above scheme further includes: The fifth information associated with the first frame stored in the database is used as the second information; the database is used to store at least one or more frames of the video corresponding to the fifth information, which is used to describe the text of the frame.
[0010] The method in the above scheme further includes: The fifth information corresponding to one or more frames of the video can be obtained through one of the following methods: The video stream of the video is acquired; key frames are extracted from the video stream to obtain one or more key frames; for the frame corresponding to each key frame, the frame is processed using OCR technology to obtain the fifth information corresponding to the frame; Obtain fifth information corresponding to one or more frames of the video sent by the video acquisition terminal; Obtain the video stream of the video and a document associated with the video, wherein the document contains text for presentation in one or more frames of the video; using the document and the video stream, determine the fifth information corresponding to one or more frames of the video.
[0011] The method in the above scheme further includes: The system receives a sixth message sent by the client, the sixth message being used to request information related to the text of the second frame of the video. If the text in the second screen is the same as the text in the first screen, a seventh message is sent to the client. The seventh message is used to instruct the client to redraw the text in the second screen using the second message corresponding to the first screen.
[0012] This application also provides an electronic device, including: a processor and a memory for storing a computer program capable of running on the processor. Wherein, when the processor is used to run the computer program, it executes the steps of any of the above-described client-side methods, or executes the steps of any of the above-described server-side methods.
[0013] This application embodiment also provides a storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of any of the above-described client-side methods or the steps of any of the above-described server-side methods.
[0014] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described client-side methods or the steps of any of the above-described server-side methods.
[0015] The information processing method, information transmission method, electronic device, storage medium, and computer program product provided in this application embodiment involve a client sending first information to a server during video playback. This first information requests information related to the text in the first frame of the video. The server receives the first information. The server then sends second information to the client, describing the text in the first frame. The client receives the second information from the server. The client then uses the second information to redraw the text in the first frame. The solution provided in this application embodiment allows the client to obtain text descriptions of the video frames from the server during video playback and redraw the text based on these descriptions to make the text in the frame clearer (or easier to recognize). This effectively improves the clarity of the text in the frame. Furthermore, transmitting text description information between the server and client for text redrawing requires fewer resources than transmitting higher-resolution video resources, better meeting practical application needs, such as ensuring the clarity of text in a weak network environment. Attached Figure Description
[0016] Figure 1 This is a schematic diagram illustrating the effect of image super-resolution (Image SR) image processing in related technologies; Figure 2 This is a schematic diagram illustrating the effect of text recognition based on OCR technology in related technologies; Figure 3 This is a schematic diagram of a video frame containing text in related technologies. Figure 4 This is a flowchart illustrating an information processing method according to an embodiment of this application; Figure 5 This is a flowchart illustrating an information transmission method according to an embodiment of this application; Figure 6 This is a flowchart illustrating a redrawing method according to an embodiment of this application; Figure 7This is a schematic diagram illustrating the process of pushing a stream from the data acquisition terminal to the server, as an application example in this application. Figure 8 This is a flowchart illustrating a live streaming method for improving image clarity, serving as an application example of this application. Figure 9 This is a schematic diagram of the structure of an information processing device according to an embodiment of this application; Figure 10 This is a schematic diagram of an information transmission device according to an embodiment of this application; Figure 11 This is a schematic diagram of the client structure in an embodiment of this application; Figure 12 This is a schematic diagram of the server structure in an embodiment of this application; Figure 13 A schematic diagram of the system structure is redrawn for the embodiments of this application. Detailed Implementation
[0017] The present application will now be described in further detail with reference to the accompanying drawings and embodiments.
[0018] Before describing the embodiments of this application, the following terms will be explained: 1) Image SR: Image SR is a computer vision technique designed to reconstruct low-resolution (LR) images into high-resolution (HR) images, effectively improving image clarity.
[0019] Currently, image processing methods using Image SR technology can be broadly categorized into three types: interpolation-based methods, reconstruction-based methods, and deep learning-based methods. Among these, deep learning-based methods may utilize algorithms such as Real-Enhanced Super-Resolution Generative Adversarial Network (Real ESRGAN), for example... Figure 1 As shown, the Real ESRGAN algorithm can effectively realize Image SR.
[0020] 2) OCR: refers to the process of converting printed, handwritten, or machine-printed text in an image into a machine-readable text format through various technologies such as pattern recognition, machine learning, and artificial intelligence.
[0021] By applying OCR technology to process images (such as video images), we can obtain information describing the text, such as the content, location information, and recognition confidence level of the text in the image.
[0022] For example, such as Figure 2 As shown, assuming the PaddleOCR tool, based on OCR technology, is used to recognize the product trademark image, it can obtain relevant information about each region containing text in the trademark image. Specifically, for the title text in the trademark, the processing result includes the following three parts: Part 1: [28.0, 37.0], [302.0, 39.0], [302.0, 72.0], [27.0, 70.0]; Part Two: XX Nourishing Hair Conditioner; Part 3: 0.9658738374710083.
[0023] The first part contains four coordinates, each corresponding to a different area of the title text (i.e., ...). Figure 2 The first part (within the bolded box) represents the four vertices of the title text. The second part represents the extracted title text content, and the third part represents the title text recognition confidence score. Recognition confidence score is an evaluation metric for the accuracy of OCR results during the OCR process. It indicates whether the text in an image is clear and legible. Recognition confidence score is typically represented as a floating-point number ranging from [0,1]. A value closer to 1 indicates a more reliable recognition result, while a value closer to 0 indicates a higher probability of recognition error.
[0024] 3) Live streaming technology: This refers to the technology of transmitting real-time audio and video content to viewers via the internet. Live streaming technology typically includes the following components: Audio and video capture: The capture end acquires (or captures) media data through external devices such as cameras and microphones; Audio and video processing (acquisition end): The acquisition end processes the media data acquired during the acquisition stage, such as performing encoding and decoding, adding special effects, etc., to obtain the processed media information; Streaming: The acquisition end transmits the processed media information to the server (which can also be understood as the server or network server) through network protocols such as Real-Time Messaging Protocol (RTMP). Audio and video processing (server): The server processes the media information pushed from the acquisition end, which may include encoding and decoding, protocol conversion, etc. Streaming refers to the process by which a client (which can also be understood as the playback end or the user's viewing end) obtains audio and video streams. Specifically, the client can obtain audio and video streams from the server through protocols such as Hypertext Transfer Protocol (HTTP) or HTTP Live Streaming (HLS); or, the client can obtain audio and video streams from the server through a Content Delivery Network (CDN). Playback: The client decodes and plays the acquired audio and video streams. At the same time, the client can selectively enhance the media data.
[0025] In related technologies, text and images in video frames are the main information mediums. When a video is played (such as live video streaming or video on demand), the client user can obtain the information carried by the video through the text and images in the frame (i.e., the video frame).
[0026] In related technologies, for the same video, the server usually supports providing video streams of multiple resolutions for the client to pull. The higher the resolution of the video stream, the better the image recognition and the more transmission resources it occupies; correspondingly, the lower the resolution of the video stream, the worse the image recognition and the less transmission resources it occupies.
[0027] Therefore, in weak network scenarios (which can also be understood as poor network conditions, network limitations, or limited transmission resources), clients typically choose to pull lower-resolution video streams from the server to reduce the resources consumed by transmission, thereby ensuring smooth video playback.
[0028] When the resolution of a video stream is low, images and text in the video frame may become difficult to recognize, failing to meet the user's viewing needs. Specifically, because text has lower recognizability and less tolerance for errors compared to images, text in a low-resolution video stream is more likely to fail to meet viewing requirements, meaning it is difficult to accurately identify.
[0029] For example, because text can carry more information than images, therefore, such as Figure 3 As shown, educational and technical videos typically contain a large amount of dense text (i.e., occupying a significant portion of the screen) to enable viewers (or users) to efficiently extract the information conveyed by the video. In this case, if the video stream retrieved by the client has low resolution, users often struggle to accurately identify the text on the screen, making it difficult to extract the information from the video.
[0030] In related technologies, the following methods can typically be used to improve text recognition: The first method: The client pulls a high-resolution video stream from the server; The second method: The client pulls multi-view video streams from the server, such as video streams for text-based perspectives; The third method is to enhance the pulled video stream locally on the client side, for example, by applying Image SR technology to process the video images.
[0031] However, for the first method, high-definition video streams may not be able to fully guarantee the clarity of the text, and they consume a lot of transmission resources; for the second method, video streams that support multiple perspectives will greatly increase costs and are difficult to apply in practice; for the third method, the application of technologies such as Image SR to process text in video images is not effective, and the processed text is often still difficult to recognize and cannot meet viewing needs.
[0032] As can be seen from the above description, there is an urgent need for a solution that can ensure the clarity of text in video footage under weak network conditions.
[0033] Based on this, in various embodiments of this application, when the client plays a video, it obtains the text description of the video frame from the server and redraws the text based on the text description to make the text in the frame clearer (or easier to recognize). In this way, on the one hand, by redrawing the text in the frame, the clarity of the text in the frame can be effectively improved; on the other hand, the transmission of text description information between the server and the client for the redrawing of the text in the frame requires fewer resources than transmitting higher-definition video resources, which can better meet the needs of practical applications, such as ensuring the clarity of the text in the frame in a weak network environment.
[0034] This application provides an information processing method applied to a client, such as... Figure 4 As shown, the method includes: Step 401: During video playback, send first information to the server, the first information being used to request information related to the text of the first frame of the video; Step 402: Receive the second information sent by the server, the second information being used to describe the text of the first screen; Step 403: Using the second information, redraw the text of the first screen.
[0035] In practical applications, the client can be referred to as a viewing end, client device, user streaming device, or user, etc. This embodiment does not limit this terminology, as long as its functionality is implemented. The server can be referred to as a server, network server, video processing server, video server, etc. This embodiment does not limit this terminology, as long as its functionality is implemented. The client is at least used to retrieve (or receive) the video stream from the server; correspondingly, the server is at least used to provide the video stream.
[0036] When the client plays a video obtained from the server (i.e., during video playback), the text on the video screen may be difficult for the user (viewer, client user, etc.) to accurately identify. In this case, the client can improve the clarity of the text on the screen by redrawing (or re-drawing). Here, redrawing refers to using a graphics library (such as OpenGL, Open Graphics Library, or Fast Forward Moving Picture Experts Group, etc.) to redraw the text on the screen clearly, obtaining text with a clarity that meets viewing requirements (this can also be understood as text areas, text portions, or text content, etc.), and then overlaying the redrawn text onto the screen for display, thereby improving the clarity and readability of the text on the screen. At the same time, the screen after text redrawing can also support operations such as copying text content, thereby improving the efficiency of users obtaining information.
[0037] In order to redraw the text on the screen, the client needs to obtain relevant information about the text in the screen. For example, the client obtains relevant information about the text in the screen from the server.
[0038] Based on this, in step 401, the client sends first information to the server for the first frame of the video (i.e., the frame where text redrawing is required), requesting information related to the first frame. Here, the first information can also be called request information, request parameters, etc., and this embodiment does not limit the name of the first information. Specifically, the first information may include the position information of the first frame in the video, such as the time elapsed since the start of video playback, so that the server can locate the first frame and determine the relevant information of the text in the first frame.
[0039] In practical applications, the client can first determine whether the clarity of the text in the first image meets the viewing requirements and obtain a judgment result. Then, if the judgment result indicates that the clarity of the text in the first image does not meet the viewing requirements, the client sends first information to the server. In this way, the client can obtain relevant information about the text in the image with low clarity, thereby further reducing the consumption of transmission resources.
[0040] Based on this, in some optional embodiments, sending the first information to the server includes: The third piece of information is determined, which indicates whether the text clarity of the first image meets the viewing requirements; If the third information indicates that the text clarity of the first image does not meet the viewing requirements, the first information is sent to the server.
[0041] In practical applications, the client can determine the third information in response to a specific user operation (which can also be understood as the user performing a specified action). The specific operation can be set according to actual needs, such as a zoom-in operation, or a click operation on a specific button (which may include a zoom button, etc.); or a click operation on the screen combined with a mouse wheel operation, etc. This embodiment does not limit the specifics. The third information can also be referred to as clarity information, etc., and this embodiment does not limit the name of the third information.
[0042] In practical applications, the client can use OCR technology to process the first screen (which can also be understood as performing OCR text detection on the first screen), obtain the processing result, and determine the third information based on the processing result. That is, in some optional embodiments, determining the third information includes: The first image is processed using OCR technology to obtain fourth information, which represents the confidence level of text recognition in the first image. The third information is determined using the fourth information and the preset threshold.
[0043] In practical applications, the fourth information is used to indicate whether the text in the first image is clear and legible. The fourth information can also be called text recognition. This application does not limit the name of the fourth information in the embodiments.
[0044] In practical applications, when the first screen contains only a single area containing text (which can also be understood as a text area or text screen, hereinafter referred to as a text area), the client can use OCR technology to process the first screen, obtain the recognition confidence of the text area in the first screen, and use the recognition confidence as the fourth information.
[0045] When there are multiple text regions in the first image, the client can process the first image based on OCR technology to obtain the recognition confidence, content and location information corresponding to each text region; then, the client can determine the fourth information by weighted summation.
[0046] Specifically, the client can determine the weight of each text region based on its location information and content, as shown in formula (1): (1) in, Indicates the first i The weight of each text region i An integer greater than or equal to 1; Indicates the first i The area of each text region can be determined by the location information corresponding to the text region. Indicates the first i The number of characters in a text region can be determined by the content corresponding to the text region; This represents a constant used for adjustment; its specific value can be set according to actual needs.
[0047] After determining the weight of each text region, the client can use the weight to determine the fourth information, which can be specifically expressed as formula (2): (2) in, The text recognition rate of the first image is represented by the fourth information. Indicates the first i The confidence level of recognition for each text region.
[0048] After determining the fourth piece of information, the client judges whether the fourth piece of information is greater than a preset threshold, thereby determining whether the text in the first image is clear enough (i.e., whether it can be accurately recognized). Specifically, when the fourth piece of information is greater than the preset threshold, the client can determine that the clarity of the first image is low and needs to obtain relevant information about the text, and then send the first piece of information to the server; when the fourth piece of information is less than or equal to the preset threshold, the client can determine that the clarity of the first image is high and can directly execute the user-requested operation, such as zooming in on the image.
[0049] Of course, in practical applications, the client can also be configured to enable the screen text redrawing function (which can also be understood as text redrawing function or redrawing function). When the client enables the screen text redrawing function, the client may not execute the above judgment process, but instead periodically send request information to the server based on a preset time interval to request the relevant information of the text on the corresponding screen.
[0050] Upon receiving the first information sent by the client, the server can provide the client with information related to the text in the first frame. Specifically, the server can transmit a portion of the frame containing the text and meeting clarity requirements (which can also be understood as a portion of the complete frame containing the text), such as capturing the text-related area from a high-definition video stream, or capturing a video stream from a multi-angle video stream at an angle related to the text. However, even if the acquired portion of the frame is smaller than the complete frame, the clarity of the portion must be high to ensure the clarity and readability of the text. Therefore, transmitting the portion still requires significant resources. In other words, the above solution is difficult to apply in weak network environments.
[0051] Based on this, the server can store text description information for each of one or more (or at least one) frames corresponding to the video, whereby the text description information describes the text in the frame. Thus, after receiving the first information, the server can filter out the second information—the text description information corresponding to the first frame—from the stored text description information; then, it sends the second information to the client for the client to redraw the text.
[0052] In practical applications, the way the text description information describes the text in the image can be set according to actual needs. For example, the text description information may include one or more of the following (one or more can also be understood as at least one): the position information of the text in the image, the content information of the text, the color information of the text, and the time information of the image corresponding to the text. Among them, the position information of the text in the image may include the coordinates of the four vertices of the text area; the content information of the text may include the text corresponding to the text (specifically, it may be encoded text or unencoded text); the color information of the text may include key-value pairs formed by hexadecimal color codes; and the time information of the image corresponding to the text may include the appearance time and disappearance time of the image containing the text.
[0053] As can be seen from the above description, the transmission of text description information between the server and the client is used for redrawing text in the image. Compared with transmitting higher-definition video resources, transmitting text description information requires fewer resources and can better meet the needs of practical applications, such as ensuring the clarity of text in the image in a weak network environment.
[0054] Accordingly, embodiments of this application also provide an information transmission method, applied to a server, such as... Figure 5 As shown, the method includes: Step 501: Receive the first information sent by the client during video playback, the first information being used to request information related to the text of the first frame of the video; Step 502: Send second information to the client. The second information is used to describe the text of the first screen and is used by the client to redraw the text of the first screen.
[0055] In practical applications, when the client plays a video obtained from the server (i.e., when playing the video), the text on the video screen may be difficult for the user (or viewer, client user, etc.) to accurately identify. In this case, the client can improve the clarity of the text on the screen by redrawing (or re-drawing).
[0056] In order to redraw the text in the first frame of the video, the client can send a first message to the server to request relevant information about the text in the first frame. Accordingly, in step 501, the server receives the first message.
[0057] To ensure text clarity in video footage under weak network conditions, the server can store text description information for each of one or more (or at least one) frames corresponding to the video. This text description information describes the text within that frame. Upon receiving the first information, the server can filter out the second information—the text description information corresponding to the first frame—from the stored text description information. Then, it sends the second information to the client for the client to redraw the text.
[0058] Based on this, in some optional embodiments, the method may further include: The fifth information associated with the first frame stored in the database is used as the second information; the database is used to store at least one or more frames of the video corresponding to the fifth information, which is used to describe the text of the frame.
[0059] In practical applications, the fifth piece of information can also be referred to as textual description information. This application embodiment does not limit the name of the fifth piece of information. The database is set on the server side (or, as can be understood, set on the server-side). The specific implementation of the database can be configured according to actual needs, and this application embodiment does not limit this.
[0060] In practical applications, before determining the second information from the stored fifth information, the server needs to first obtain the fifth information corresponding to one or more frames of the video and store the obtained fifth information in the database.
[0061] Specifically, in some optional embodiments, the method may further include: The fifth information corresponding to one or more frames of the video can be obtained through one of the following methods: The first method is to acquire the video stream of the video; extract keyframes from the video stream to obtain one or more keyframes; and process the image corresponding to each keyframe using OCR technology to obtain the fifth information corresponding to the image. The second method: Obtain the fifth information corresponding to one or more frames of the video sent by the video acquisition terminal; The third method involves obtaining the video stream of the video and a document associated with the video, wherein the text contained in the document is used to be presented in one or more frames of the video; and using the document and the video stream, determining the fifth information corresponding to one or more frames of the video.
[0062] For the first method described above, the specific process may include: the server acquiring the video stream from the acquisition end (which can also be understood as the video provider or uploader, etc.); then, the server extracting (or reading) keyframes from the video; for each keyframe, the server using OCR technology to identify the region information and content information of the text corresponding to that frame, and calling relevant tools (which can be selected according to actual needs) to obtain the time information and color information of the text corresponding to that frame; finally, all the information obtained from the text used to describe the frame (such as region information, content information, time information, color information, etc.) is used as the fifth piece of information for that frame. Here, a keyframe refers to a representative frame in the video stream, which usually contains a lot of visual information and is commonly used in scenarios such as video content summarization, indexing, and content analysis.
[0063] Specifically, after obtaining the fifth information corresponding to each keyframe, the server can perform deduplication processing on the fifth information corresponding to adjacent keyframes, storing only the deduplicated fifth information, thereby avoiding the storage of a large amount of identical fifth information and reducing the storage resources occupied. Alternatively, after extracting multiple keyframes, the server can first perform deduplication processing on the images corresponding to the extracted keyframes (e.g., deduplication based on inter-frame difference), and then determine the corresponding fifth information for the images of the deduplicated keyframes. The server can choose the actual deduplication method according to actual needs (e.g., performance requirements, resource requirements), and this embodiment does not limit this approach.
[0064] Regarding the second method described above, the video acquisition end can extract keyframes from the video locally and use OCR technology and related tools to determine the fifth information corresponding to each keyframe (that is, moving the process of obtaining the fifth information forward to the video acquisition stage). In this way, the acquisition end can directly send the determined fifth information to the server; correspondingly, the server can directly obtain the fifth information of the frame corresponding to each keyframe, which can effectively reduce the resource consumption of the server.
[0065] Regarding the third method described above, the server can obtain the video stream and the document associated with the video from the video acquisition end. The document associated with the video may include presentations (such as PPT), text documents (such as PDF, Word, etc.), subtitle files, etc., for display in the video. This application embodiment does not limit this.
[0066] After obtaining the document and video stream, the server can determine the fifth information corresponding to each frame in one or more frames of the video stream, based on the content in the document that matches the content displayed in that frame.
[0067] In practical applications, the server can use one of the three methods mentioned above to obtain the fifth information corresponding to one or more frames of the video, as needed; then, the server can store all the obtained fifth information.
[0068] In practical applications, the client may send multiple requests to the server to request text information for multiple frames. For example, after sending the first message, the client can send a sixth message to the server to request text information for a second frame, where the second frame refers to the first frame in the video that requires text redrawing after the first frame (or the frame following the first frame). In this case, if the text in the second frame is unchanged from the first frame, the server does not need to repeatedly send the text description information of the frame to the client.
[0069] Based on this, in some optional embodiments, the method may further include: The system receives a sixth message sent by the client, the sixth message being used to request information related to the text of the second frame of the video. If the text in the second screen is the same as the text in the first screen, a seventh message is sent to the client. The seventh message is used to instruct the client to redraw the text in the second screen using the second message corresponding to the first screen.
[0070] In practical applications, the seventh piece of information can also be understood as an agreed-upon unchanging identifier (or mark), such as a repeat identifier, etc. This application embodiment does not limit this.
[0071] This application also provides a screen redrawing method, such as... Figure 6 As shown, the method includes: Step 601: During video playback, the client sends first information to the server, the first information being used to request information related to the text of the first frame of the video; Step 602: The server receives the first information sent by the client; Step 603: The server sends second information to the client, the second information being used to describe the text of the first screen; Step 604: The client receives the second information sent by the server; Step 605: The client uses the second information to redraw the text on the first screen.
[0072] It should be noted that the specific processing procedures of the client and server have been described in detail above, and will not be repeated here.
[0073] The measurement configuration method provided in this application embodiment involves a client sending first information to a server during video playback. This first information requests information related to the text in the first frame of the video. The server receives the first information. The server then sends second information to the client, describing the text in the first frame. The client receives the second information from the server. The client uses the second information to redraw the text in the first frame. The solution provided in this application embodiment allows the client to obtain text descriptions of the video frames from the server during video playback and redraw the text based on these descriptions to make the text in the frame clearer (or easier to recognize). Thus, on the one hand, redrawing the text in the frame effectively improves the clarity of the text; on the other hand, transmitting text description information between the server and client for text redrawing requires fewer resources than transmitting higher-resolution video resources, better meeting practical application needs, such as ensuring the clarity of text in a weak network environment.
[0074] The following section provides a more detailed description of this application with reference to application examples.
[0075] This application example provides a live streaming system for improving image clarity, which may include a capture terminal, a server terminal, and a client terminal. Among them, such as... Figure 7 As shown, the acquisition end (also known as the video acquisition end, such as a camera) is used to capture the broadcaster's video footage, and after processing the captured video footage, it is pushed to the streaming media server and video processing server on the server side; the streaming media server is used to provide live broadcast-related services, including protocol conversion, mixing, slicing, etc.; the video processing server reads keyframes in the video stream, recognizes text images and text positions through OCR technology, removes duplicate text images from keyframes, records the time, area, color, content, etc. of the text images, as shown in Table 1, and saves them to the database;
[0076] Table 1. Data saved after image recognition by the video processing service. Of course, text recognition, deduplication, and database entry requests can also be performed by the video capture device. After receiving the database entry request from the video capture device, the server stores the relevant information of the text image sent by the video capture device into the database. The video at the video capture end is the original video stream with the highest clarity. By using OCR technology to recognize text at the capture end, the best recognition effect can be obtained.
[0077] Based on the above system, this application provides a live streaming method to improve image clarity, such as... Figure 8As shown, it includes the following steps: Step 801: After the client (specifically, the player in the client) obtains the user's (or the live stream viewer's) selection action on the live stream topic (such as clicking on the live stream content list or clicking to enter the live stream room), it pulls the stream for playback; then, it executes step 802. Step 802: The client determines whether the user has triggered a click to zoom in operation and obtains the result; In practical applications, the user's specified behavior can be pre-defined as wanting to zoom in on the screen and written into the player's function description, such as clicking a specific button (like a magnifying glass) or clicking on the video screen and scrolling the mouse wheel.
[0078] If the judgment result indicates that a click-to-zoom operation is triggered, proceed to step 803; or, if the judgment result indicates that a click-to-zoom operation is not triggered, proceed to step 804.
[0079] Step 803: The client detects the text recognition accuracy of the current screen and obtains the detection result; If the detection result indicates that the text is clear, proceed directly to zoom in on the image; or, if the detection result indicates that the text is unclear, proceed to step 805.
[0080] Step 804: The client determines whether the text redraw function is enabled and obtains the result; If the judgment result indicates that the text redrawing function is not enabled, continue streaming playback; or, When the judgment result indicates that the text redrawing function is enabled, step 805 is executed.
[0081] Step 805: The client queries the server for the text information of the current screen (such as the first screen mentioned above) (i.e., the fifth information corresponding to the screen) and obtains the query result; Specifically, when a client queries the server for the text content of the current live stream, it needs to provide the server with request parameters, such as the current position of the live stream within the live video (e.g., how many milliseconds have passed since the start of the live stream, which can also be understood as a position timestamp (ts)). This allows the server to determine the original text information content identified and saved by the video server or video capture device in the server-side processing flow based on this positional information. Simultaneously, if the client is not sending a request (i.e., performing a query) for the first time, the client can also provide the server with the timestamp (prev_ts) of the previous request, allowing the server to determine whether the text content in the video has changed relative to the previous request.
[0082] For example, based on the content shown in Table 1, assuming that the client performs a query at 600001 milliseconds, since the text content has not changed before 615000 milliseconds, that is, the data is the same, the server does not need to reply to the client with the text content again when other query requests are received before 615000 milliseconds.
[0083] After obtaining the query results, the client uses the query results to determine whether the text on the current screen has been updated, and obtains a judgment result; if the judgment result indicates that there is an update, step 806 is executed; or, if the judgment result indicates that there is no update, the streaming playback continues.
[0084] Step 806: The client redraws the text on the current screen and performs super-resolution processing on the image of the current screen.
[0085] Specifically, the client can improve the clarity of text images by redrawing and overlaying the original image. At the same time, the client can perform Image SR reconstruction algorithm on non-text parts of the image to improve the clarity of non-text parts.
[0086] Of course, the client can also display the redrawn text in other ways, such as by providing recognition and parsing of the text on the screen through a pop-up window (i.e., displaying it in boxes).
[0087] In practical applications, when the video frame includes a PowerPoint presentation, all areas outside the PowerPoint can be considered non-text areas. In this case, the client can distinguish between text and non-text areas using text area information obtained through OCR technology (such as the vertex coordinates of a rectangle). Furthermore, since text in a PowerPoint presentation is relatively regular and dense, while text in a video is often distorted, the client can further distinguish between text and non-text areas using text density (determined by the text content and rectangle size) and the shape of the text area rectangle (determined by the vertex coordinates of the rectangle).
[0088] In practical applications, the client can use other video processing methods such as inter-frame difference to detect continuously changing regions and discontinuously changing regions (such as text regions in PPT). Then, Image SR processing is performed on continuously changing regions, and redrawing and overlay enhancement is performed on text regions.
[0089] The solution provided in this application example primarily focuses on the audio and video processing (capture end, server end) and playback stages. It processes audio and video into multiple resolutions, identifies and stores text within the video frame, reads the text during playback, and redraws the text in the player. This additional text information improves the clarity of the text-based visuals on the user's viewing end. Simultaneously, it applies image super-resolution (SR) technology to non-text-based visuals on the user's viewing end to enhance their clarity, thereby improving the efficiency of live information delivery. Furthermore, the text-based visuals allow users to further process information, enhancing the user experience.
[0090] In this way, under weak network conditions, the client will adaptively pull data from a lower-resolution data source. The solution provided in this application example recognizes text in the high-definition video at the acquisition end and transmits the text to the client via the server when needed. This method uses text instead of image text, resulting in low storage and transmission costs. It enables high-definition live streaming even in poor network conditions, avoiding information loss caused by blurry text and improving the viewing experience under weak network conditions. Thus, the problem of unrecognizable text due to blurry images can be corrected, and text information in the live stream can be preserved to the greatest extent, thereby compensating for information loss caused by blurry images under weak network conditions and achieving high-definition live streaming under these conditions. Of course, the solution provided in this application example can also be applied to the video-on-demand field.
[0091] To implement the method provided on the client side in this application embodiment, this application embodiment also provides an information processing device, which is set on the client, such as... Figure 9 As shown, the device includes: The first sending unit 901 is used to send first information to the server during video playback, wherein the first information is used to request the acquisition of relevant information of the text of the first frame of the video. The first receiving unit 902 is used to receive second information sent by the server, the second information being used to describe the text of the first screen; The redrawing unit 903 is used to redraw the text of the first screen using the second information.
[0092] In some optional embodiments, the first transmitting unit 901 is specifically used for: The third piece of information is determined, which indicates whether the text clarity of the first image meets the viewing requirements; If the third information indicates that the text clarity of the first image does not meet the viewing requirements, the first information is sent to the server.
[0093] In some optional embodiments, the first transmitting unit 901 is specifically used for: The first image is processed using OCR technology to obtain fourth information, which represents the confidence level of text recognition in the first image. The third information is determined using the fourth information and the preset threshold.
[0094] In practical applications, the first receiving unit 902 can be implemented by the communication interface in the information processing device, the first sending unit 901 can be implemented by the processor in the information processing device in combination with the communication interface, and the redrawing unit 903 can be implemented by the processor in the information processing device.
[0095] It should be noted that the information processing apparatus provided in the above embodiments is only illustrated by the division of the above-described program units. In practical applications, the above processing can be assigned to different program units as needed, that is, the internal structure of the apparatus can be divided into different program units to complete all or part of the processing described above. In addition, the information processing apparatus and information processing method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
[0096] To implement the server-side method of this application embodiment, this application embodiment also provides an information transmission device, which is set on the server, such as... Figure 10 As shown, the device includes: The second receiving unit 1001 is used to receive first information sent by the client during video playback, the first information being used to request information related to the text of the first frame of the video; The second sending unit 1002 is used to send second information to the client. The second information is used to describe the text of the first screen and is used by the client to redraw the text of the first screen.
[0097] In some optional embodiments, the device may further include: The processing unit is configured to use the fifth information associated with the first frame stored in the database as the second information; the database is configured to store at least one or more frames of the video corresponding to the fifth information, the fifth information being used to describe the frame.
[0098] In some optional embodiments, the processing unit is further configured to: The fifth information corresponding to one or more frames of the video can be obtained through one of the following methods: The video stream of the video is acquired; key frames are extracted from the video stream to obtain one or more key frames; for the frame corresponding to each key frame, the frame is processed using OCR technology to obtain the fifth information corresponding to the frame; Obtain fifth information corresponding to one or more frames of the video sent by the video acquisition terminal; Obtain the video stream of the video and a document associated with the video, wherein the document contains text for presentation in one or more frames of the video; using the document and the video stream, determine the fifth information corresponding to one or more frames of the video.
[0099] In some optional embodiments, the second receiving unit 1001 is further configured to: The system receives a sixth message sent by the client, the sixth message being used to request information related to the text of the second frame of the video. The second transmitting unit 1002 is further configured to: If the text in the second screen is the same as the text in the first screen, a seventh message is sent to the client. The seventh message is used to instruct the client to redraw the text in the second screen using the second message corresponding to the first screen.
[0100] In practical applications, the second receiving unit 1001 and the second sending unit 1002 can be implemented by the communication interface in the information transmission device, and the processing unit can be implemented by the processor in the information transmission device.
[0101] It should be noted that the information transmission device provided in the above embodiments is only illustrated by the division of the above-described program units. In practical applications, the above processing can be assigned to different program units as needed, that is, the internal structure of the device can be divided into different program units to complete all or part of the processing described above. In addition, the information transmission device and the information transmission method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
[0102] Based on the hardware implementation of the above program modules, and in order to implement the client-side method of this application embodiment, this application embodiment also provides a client, such as... Figure 11 As shown, the client 1100 includes: The first communication interface 1101 is capable of exchanging information with the server; The first processor 1102 is connected to the first communication interface 1101 to enable information interaction with the server and to execute the methods provided by one or more of the above-mentioned client-side technical solutions when running a computer program. The computer program is stored in the first memory 1103.
[0103] Specifically, the first communication interface 1101 is used for: During video playback, a first message is sent to the server, the first message being used to request information related to the text of the first frame of the video; and a second message is received from the server, the second message being used to describe the text of the first frame. The first processor 1102 is used for: Using the second information, redraw the text in the first image.
[0104] In some optional embodiments, the first processor 1102 is specifically used for: The third piece of information is determined, which indicates whether the text clarity of the first image meets the viewing requirements; If the clarity of the text in the first image does not meet the viewing requirements as indicated by the third information, the first information is sent to the server through the first communication interface 1101.
[0105] In some optional embodiments, the first processor 1102 is specifically used for: The first image is processed using OCR technology to obtain fourth information, which represents the confidence level of text recognition in the first image. The third information is determined using the fourth information and the preset threshold.
[0106] It should be noted that the specific processing procedures of the first processor 1102 and the first communication interface 1101 can be understood by referring to the above method.
[0107] Of course, in practical applications, the various components in client 1100 are coupled together through bus system 1104. It can be understood that bus system 1104 is used to implement communication between these components. In addition to a data bus, bus system 1104 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 11 The general designated all buses as Bus System 1104.
[0108] The first memory 1103 in this embodiment is used to store various types of data to support the operation of the client 1100. Examples of such data include any computer program used to operate on the client 1100.
[0109] The methods disclosed in the above embodiments of this application can be applied to the first processor 1102, or implemented by the first processor 1102. The first processor 1102 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware or by instructions in the form of software in the first processor 1102. The first processor 1102 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The first processor 1102 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in the first memory 1103. The first processor 1102 reads the information in the first memory 1103 and completes the steps of the aforementioned method in combination with its hardware.
[0110] In an exemplary embodiment, the client 1100 may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the aforementioned method.
[0111] Based on the hardware implementation of the above program modules, and in order to implement the server-side method of this application embodiment, this application embodiment also provides a server, such as... Figure 12 As shown, the server 1200 includes: The second communication interface 1201 is capable of exchanging information with the client; The second processor 1202 is connected to the second communication interface 1201 to enable information interaction with the client and to execute the methods provided by one or more of the above-mentioned server-side technical solutions when running computer programs. The computer program is stored in the second memory 1203.
[0112] Specifically, the second communication interface 1201 is used for: Receive first information sent by the client during video playback, the first information being used to request relevant information about the text of the first frame of the video; Send a second message to the client. The second message is used to describe the text of the first screen. The second message is used to allow the client to redraw the text of the first screen.
[0113] In some optional embodiments, the second processor 1202 is used for: The fifth information associated with the first frame stored in the database is used as the second information; the database is used to store at least one or more frames of the video corresponding to the fifth information, which is used to describe the text of the frame.
[0114] In some optional embodiments, the second processor 1202 is further configured to: The fifth information corresponding to one or more frames of the video can be obtained through one of the following methods: The video stream of the video is acquired; key frames are extracted from the video stream to obtain one or more key frames; for the frame corresponding to each key frame, the frame is processed using OCR technology to obtain the fifth information corresponding to the frame; Obtain fifth information corresponding to one or more frames of the video sent by the video acquisition terminal; Obtain the video stream of the video and a document associated with the video, wherein the document contains text for presentation in one or more frames of the video; using the document and the video stream, determine the fifth information corresponding to one or more frames of the video.
[0115] In some optional embodiments, the second communication interface 1201 is further used for: The system receives a sixth message sent by the client, the sixth message being used to request information related to the text of the second frame of the video. If the text in the second screen is the same as the text in the first screen, a seventh message is sent to the client. The seventh message is used to instruct the client to redraw the text in the second screen using the second message corresponding to the first screen.
[0116] It should be noted that the specific processing procedures of the second processor 1202 and the second communication interface 1201 can be understood by referring to the above method.
[0117] Of course, in practical applications, the various components in server 1200 are coupled together through bus system 1204. It can be understood that bus system 1204 is used to implement communication between these components. In addition to a data bus, bus system 1204 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 12 The general labeled all buses as Bus System 1204.
[0118] The second memory 1203 in this embodiment is used to store various types of data to support the operation of the server 1200. Examples of such data include any computer program used to operate on the server 1200.
[0119] The methods disclosed in the embodiments of this application can be applied to the second processor 1202, or implemented by the second processor 1202. The second processor 1202 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware or by instructions in the form of software in the second processor 1202. The second processor 1202 may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The second processor 1202 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in the second memory 1203. The second processor 1202 reads the information in the second memory 1203 and completes the steps of the aforementioned method in combination with its hardware.
[0120] In an exemplary embodiment, server 1200 may be implemented by one or more ASICs, DSPs, PLDs, CPLDs, FPGAs, general-purpose processors, controllers, MCUs, microprocessors, or other electronic components to perform the aforementioned method.
[0121] It is understood that the memories (first memory 1103, second memory 1203) in the embodiments of this application can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM).The memories described in the embodiments of this application are intended to include, but are not limited to, these and any other suitable types of memories.
[0122] In an exemplary embodiment, this application also provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium. For example, it may include a first memory 1103 storing a computer program, which can be executed by a first processor 1102 of a client 1100 to complete the steps described in the aforementioned client-side method. Another example is a second memory 1203 storing a computer program, which can be executed by a second processor 1202 of a server 1200 to complete the steps described in the aforementioned server-side method. The computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM.
[0123] In an exemplary embodiment, this application also provides a computer program product, including a computer program that can be executed by a first processor 1102 of a client 1100 to complete the steps described in the aforementioned client-side method, or the computer program can be executed by a second processor 1202 of a server 1200 to complete the steps described in the aforementioned server-side method.
[0124] To implement the method provided in the embodiments of this application, the embodiments of this application also provide a screen redrawing system, such as... Figure 13 As shown, the system includes: client 1301 and server 1302.
[0125] It should be noted that the specific processing procedures of the client 1301 and the server 1302 have been described in detail above and will not be repeated here.
[0126] It should be noted that terms such as "first" and "second" are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0127] Furthermore, the technical solutions described in the embodiments of this application can be combined arbitrarily without conflict.
[0128] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application.
Claims
1. An information processing method, characterized in that, Applied to the client side, including: During video playback, a first message is sent to the server, the first message being used to request information related to the text of the first frame of the video; Receive second information sent by the server, the second information being used to describe the text of the first screen; Using the second information, redraw the text in the first image.
2. The method according to claim 1, characterized in that, Sending the first information to the server includes: The third piece of information is determined, which indicates whether the text clarity of the first image meets the viewing requirements; If the third information indicates that the text clarity of the first image does not meet the viewing requirements, the first information is sent to the server.
3. The method according to claim 2, characterized in that, The determination of the third information includes: The first image is processed based on optical character recognition (OCR) technology to obtain fourth information, which represents the recognition confidence of the text in the first image. The third information is determined using the fourth information and the preset threshold.
4. An information transmission method, characterized in that, Applied to servers, including: Receive first information sent by the client during video playback, the first information being used to request relevant information about the text of the first frame of the video; Send a second message to the client. The second message is used to describe the text of the first screen. The second message is used to allow the client to redraw the text of the first screen.
5. The method according to claim 4, characterized in that, The method further includes: The fifth information associated with the first frame stored in the database is used as the second information; the database is used to store at least one or more frames of the video corresponding to the fifth information, which is used to describe the text of the frame.
6. The method according to claim 5, characterized in that, The method further includes: The fifth information corresponding to one or more frames of the video can be obtained through one of the following methods: The video stream of the video is acquired; key frames are extracted from the video stream to obtain one or more key frames; for the frame corresponding to each key frame, the frame is processed using OCR technology to obtain the fifth information corresponding to the frame; Obtain fifth information corresponding to one or more frames of the video sent by the video acquisition terminal; Obtain the video stream of the video and a document associated with the video, wherein the document contains text for presentation in one or more frames of the video; using the document and the video stream, determine the fifth information corresponding to one or more frames of the video.
7. The method according to any one of claims 4 to 6, characterized in that, The method further includes: The system receives a sixth message sent by the client, the sixth message being used to request information related to the text of the second frame of the video. If the text in the second screen is the same as the text in the first screen, a seventh message is sent to the client. The seventh message is used to instruct the client to redraw the text in the second screen using the second message corresponding to the first screen.
8. An electronic device, characterized in that, include: The processor and the memory used to store computer programs that can run on the processor. When the processor is used to run the computer program, it performs the steps of the method according to any one of claims 1 to 3, or performs the steps of the method according to any one of claims 4 to 7.
9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 3, or the steps of the method according to any one of claims 4 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 3, or the steps of the method according to any one of claims 4 to 7.