Display devices and media synthesis methods
By processing the original image, audio, and text data through the controller of the display device, the problems of slow video compositing speed and low richness of the display device are solved, and efficient media asset compositing is achieved.
Patent Information
- Application Number
- CN202411245494.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-05
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2044-09-05
AI Technical Summary
Existing display devices are slow in video editing and compositing processes and have limited raw materials, resulting in low richness in the produced videos.
The controller of the display device extracts the original image data into regions in pixels, encodes the video frame data according to the preset encoding format, performs format conversion processing on the original audio data, adds time stamps to the original text data, and performs encapsulation processing on the video data, audio data and subtitle data to generate composite media asset data.
It improves the speed and richness of media asset synthesis, and can process various types of raw media asset data in real time to generate high-quality synthesized media asset data.
Smart Images

Figure CN119094808B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of display devices, and in particular to a display device and a media asset synthesis method. BACKGROUND
[0002] A display device is an intelligent device capable of presenting a user interface and supporting user interaction. Taking a smart television as an example, a smart television is a television product based on Internet application technology, having an open operating system and a chip, and possessing an open application platform, and capable of realizing bidirectional man-machine interaction functions, and integrating audio, video, entertainment, data and other functions, for meeting the diversified and personalized needs of users. A display device can play videos of various sources, including videos in an online video platform, videos in a live broadcast software and videos made by a user.
[0003] A user can make a video through a specific application in the display device. For example, a video that is shot in advance is uploaded to the display device, and an application program on the display device is used to perform editing operations such as cropping and splicing on the video; or an existing video on the display device is selected to perform a synthesis operation.
[0004] However, the method of editing a shot video or synthesizing multiple videos to make a video is slow, and the original materials required are limited, resulting in low richness of the made video. SUMMARY
[0005] The present application provides a display device and a media asset synthesis method to solve the problems of slow speed and low richness of media asset data synthesis.
[0006] In a first aspect, some embodiments of the present application provide a display device, comprising a display and a controller.
[0007] The display is configured to display a user interface; and the controller is configured to:
[0008] in response to a media asset synthesis instruction input by a user, acquire original media asset data, the original media asset data comprising original picture data, original audio data and original text data;
[0009] perform region-wise extraction on the original picture data in units of pixels to obtain video frame data;
[0010] perform encoding on the video frame data in a preset encoding format to obtain video data;
[0011] perform format conversion processing on the original audio data to obtain audio data in a preset encoding format;
[0012] add a time mark to the original text data to obtain subtitle data;
[0013] performing packaging processing on the video data, the audio data and the subtitle data to obtain synthesized media data, wherein a playing time length of the synthesized media data is equal to a playing time length of the original audio data.
[0014] In some embodiments of the present application, before the step of performing the sub-region extraction on the original picture data in units of pixels to obtain the video frame data, the controller is further configured to perform decoding on the original picture data to obtain decoded picture data, and perform data format conversion on the decoded picture data to obtain decoded picture data in a preset format if the data format of the decoded picture data is different from the preset format.
[0015] In some embodiments of the present application, the controller performing the sub-region extraction on the original picture data in units of pixels to obtain the video frame data is specifically configured to: obtain a first boundary point and a second boundary point of the decoded picture data, wherein the first boundary point and the second boundary point are two symmetrical vertices in the decoded picture data; set an extraction window according to the first boundary point and the second boundary point, wherein the width of the extraction window is equal to a preset proportion of the pixel points contained between the first boundary point and the second boundary point; and extract pixel points located in the extraction window in the decoded picture data in a first direction in sequence, taking the first boundary point as the starting point of the extraction window and taking a preset number as the translation step of the extraction window, to obtain a plurality of the video frame data, wherein the first direction is from the first boundary point to the second boundary point, and the video frame data includes an extraction boundary point.
[0016] In some embodiments of the present application, the controller performing the sub-region extraction on the original picture data in units of pixels to obtain the video frame data is further configured to: if the extraction boundary point of the video frame data coincides with the second boundary point, then extract pixel points located in the extraction window in the decoded picture data in a second direction in sequence, taking the second boundary point as the starting point of the extraction window and taking the preset number as the translation step of the extraction window, to obtain a plurality of the video frame data, wherein the second direction is from the second boundary point to the first boundary point.
[0017] In some embodiments of the present application, the controller performs a format conversion process on the original audio data to obtain audio data in a preset encoding format, and is specifically configured to: acquire new audio data; decodes the original audio data to obtain first decoded audio data; decodes the new audio data to obtain second decoded audio data; synthesizes the first decoded audio data and the second decoded audio data to obtain synthesized audio data; and encodes the synthesized audio data to obtain audio data in a preset encoding format.
[0018] In some embodiments of the present application, the controller adds time markers to the original text data to obtain subtitle data, and is specifically configured to: acquire a sampling frequency and a sampling point number of the original audio data; calculates a start time point and an end time point of the original audio data according to the sampling frequency and the sampling point number; and adds the start time point and the end time point of the original audio data to the original text data to obtain the subtitle data.
[0019] In some embodiments of the present application, the controller performs an encapsulation process on the video data, the audio data and the subtitle data to obtain synthesized media data, and is specifically configured to: generate an initial media template, the initial media template including a preset media format; input the video data, the audio data and the subtitle data into the initial media template; acquire an end identifier; and generate the synthesized media data in the preset media format according to the initial media template based on the end identifier.
[0020] In some embodiments of the present application, after the step of inputting the video data, the audio data and the subtitle data into the initial media template, the controller is further configured to: acquire a start time marker of the subtitle data; and add the subtitle data in the initial media template every preset time interval starting from a time point represented by the start time marker.
[0021] In some embodiments of the present application, after the step of acquiring original media data in response to a media synthesis instruction input by a user, the controller is further configured to: acquire a data length, a sampling frequency, a channel number and a bit depth of the original audio data; and calculate a playing duration of the original audio data according to the data length, the sampling frequency, the channel number and the bit depth.
[0022] In a second aspect, some embodiments of the present application further provide a media synthesis method applied to the display device provided in the first aspect, and the display device includes a display and a controller. The method includes:
[0023] In response to a media synthesis instruction input by a user, acquiring original media data, the original media data including original picture data, original audio data and original text data;
[0024] performing sub-region extraction on the original picture data in units of pixels to obtain video frame data;
[0025] performing encoding on the video frame data according to a preset encoding format to obtain video data;
[0026] performing format conversion processing on the original audio data to obtain audio data in a preset encoding format;
[0027] adding time markers to the original text data to obtain subtitle data;
[0028] performing encapsulation processing on the video data, the audio data, and the subtitle data to obtain synthesized media data; wherein a playing time length of the synthesized media data is equal to a playing time length of the original audio data.
[0029] According to the technical solution, some embodiments of the present application provide a display device and a media synthesis method. The method can acquire original media data in response to a media synthesis instruction input by a user; perform sub-region extraction on original picture data in units of pixels to obtain video frame data; perform encoding on the video frame data according to a preset encoding format to obtain video data; perform format conversion processing on the original audio data to obtain audio data in a preset encoding format; add time markers to the original text data to obtain subtitle data; and perform encapsulation processing on the video data, the audio data, and the subtitle data to obtain synthesized media data. The method can acquire original media data of various types, perform corresponding processing on the acquired original media data in real time, and then perform encapsulation on the processed media data to obtain synthesized media data, thereby improving the speed and richness of the synthesized media data. BRIEF DESCRIPTION OF DRAWINGS
[0030] To make the technical solutions of the present application clearer, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, other drawings can also be obtained by those skilled in the art without any creative effort on the basis of these drawings.
[0031] Figure 1 a schematic diagram of an operation scenario between a display device and a control device according to some embodiments of the present application;
[0032] Figure 2 a schematic diagram of a hardware configuration of a display device according to some embodiments of the present application;
[0033] Figure 3 a schematic diagram of a software configuration of a display device according to some embodiments of the present application;
[0034] Figure 4 A method for a controller to control a media synthesis application to synthesize media data is shown in the flowchart of FIG. 7 for some embodiments of the application;
[0035] Figure 5 A method for a media synthesis application to synthesize media data is shown in the flowchart of FIG. 8 for some embodiments of the application;
[0036] Figure 6 A method for a media synthesis application to synthesize media data is shown in the timing diagram of FIG. 9 for some embodiments of the application;
[0037] Figure 7 A method for a media synthesis application to synthesize media data is shown in the block diagram of FIG. 10 for some embodiments of the application;
[0038] Figure 8 A method for a controller to perform region-wise extraction is shown in the flowchart of FIG. 11 for some embodiments of the application;
[0039] Figure 9 A first picture obtained by a controller is shown in the flowchart of FIG. 12 for some embodiments of the application;
[0040] Figure 10 An extraction window determined by a controller is shown in the flowchart of FIG. 13 for some embodiments of the application;
[0041] Figure 11 A second stage is shown in the flowchart of FIG. 14 for some embodiments of the application;
[0042] Figure 12 A method for a controller to obtain video data is shown in the flowchart of FIG. 15 for some embodiments of the application;
[0043] Figure 13 A method for a controller to synthesize audio data is shown in the flowchart of FIG. 16 for some embodiments of the application;
[0044] Figure 14 A method for a controller to generate subtitle data is shown in the flowchart of FIG. 17 for some embodiments of the application;
[0045] Figure 15 A method for a controller to generate synthesized media data is shown in the flowchart of FIG. 18 for some embodiments of the application. DETAILED DESCRIPTION
[0046] Embodiments will be described in detail below with reference to the drawings. Descriptions of well-known functions and structures are omitted so as not to obscure the disclosure with unnecessary detail. Like reference numerals refer to like elements throughout. The descriptions of embodiments in this detailed description do not represent all of the techniques consistent with the present application. Instead, they are simply some of the techniques consistent with the present application.
[0047] It should be noted that the brief description of the terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of the application. Unless otherwise specified, these terms should be understood in accordance with their ordinary and general meanings.
[0048] The terms "first", "second", "third" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar or like objects or entities, and do not necessarily mean a specific order or sequence, unless otherwise noted. It should be understood that the terms used in this way can be interchanged under appropriate circumstances.
[0049] The terms "include" and "have" and any variations thereof are intended to cover but not exclusive inclusion, for example, a product or device including a series of components does not have to be limited to all components clearly listed, but can include other components not clearly listed or inherent to these products or devices.
[0050] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or combination of hardware or / and software code capable of performing functions associated with the element.
[0051] In the embodiments of the present application, the display device 200 generally refers to a device with picture display and data processing capability. For example, the display device 200 includes but is not limited to smart TV, mobile terminal, computer, monitor, advertising screen, wearable device, virtual reality device, augmented reality device, etc.
[0052] Figure 1 The schematic diagram of the operation scenario between the display device and the control device provided by some embodiments of the present application is shown. As shown in Figure 1 The user can operate the display device 200 through touch operation, mobile terminal 300 and control device 100. Among them, the control device 100 is used to receive the operation instruction input by the user, and convert the operation instruction into control instruction which can be recognized and responded by the display device 200. For example, the control device 100 can be a remote controller, a stylus, a handle, etc.
[0053] The mobile terminal 300 can be used as a kind of control device to perform human-computer interaction between the user and the display device 200. The mobile terminal 300 can also be used as a kind of communication device to establish communication connection with the display device 200 and perform data interaction. In some embodiments, the mobile terminal 300 can install software application with the display device 200, realize connection communication through network communication protocol, and achieve the purpose of one-to-one control operation and data communication. The mobile terminal 300 can also display audio and video content on the display device 200 to realize synchronous display function.
[0054] In some embodiments, the mobile terminal 300 or other electronic device can also simulate the functions of the control device 100 by running an application program that controls the display device 200.
[0055] As shown in Figure 1 Further shown in
[0056] The display device 200 can provide a broadcast receiving television function, and can also provide an intelligent network television function with computer support, including but not limited to, network television, smart television, Internet protocol television (IPTV), etc.
[0057] Figure 2 The hardware configuration of the display device 200 is shown in Figure 1 The hardware configuration of the display device 200 is shown in
[0058] In some embodiments, the display device 200 can include at least one of a tuning demodulator 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface 280.
[0059] In some embodiments, the detector 230 is configured to collect signals from the external environment or interaction. For example, the detector 230 can include a light receiver for collecting ambient light intensity, or an image collector such as a camera for collecting external environmental scenes, user attributes or user interaction gestures, or a sound collector such as a microphone for receiving external sounds.
[0060] In some embodiments, the display 260 includes a display function component for presenting a picture, and a driving component for driving image display. The display 260 is configured to receive image signals from the controller 250 for display. For example, the display 260 can be configured to display video content, image content, and components of a menu control interface, as well as a user control UI interface, etc.
[0061] In some embodiments, the communication device 220 is a component for communicating with the external device or the server 400 according to various communication protocol types. The display device 200 can be provided with multiple communication devices 220 according to different supported communication manners. For example, when the display device 200 supports wireless network communication, the display device 200 can be provided with a communication device 220 containing WiFi function. When the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 containing Bluetooth function.
[0062] The communication device 220 can connect the display device 200 with the external device or the server 400 through wireless or wired connection. The wired connection can connect the display device 200 with the external device through data line, interface, etc. The wireless connection can connect the display device 200 with the external device through wireless signal or wireless network. The display device 200 can directly establish connection relationship with the external device, or indirectly establish connection relationship through gateway, route, connection device, etc.
[0063] In some embodiments, the controller 250 can include at least one of central processor, video processor, audio processor, graphics processor, power supply processor, first interface to nth interface for input / output, and the controller 250 controls the operation of the display device and responds to the user's operation through various software control programs stored on the memory. The controller 250 controls the overall operation of the display device 200.
[0064] In some embodiments, the controller 250 and the tuner demodulator 210 can be located in different split devices, i.e. the tuner demodulator 210 can also be in the external device of the main device where the controller 250 is located, such as external set-top box, etc.
[0065] In some embodiments, the user can input user commands through the graphical user interface (GUI) displayed on the display 260, and the user input interface receives the user input commands through the graphical user interface (GUI).
[0066] In some embodiments, the audio output device 270 can be the native loudspeaker of the display device 200, or can be the audio output device connected to the display device 200. For the audio output device connected to the display device 200, the display device 200 can also be provided with an external audio output terminal, and the audio output device can be connected to the display device 200 through the external audio output terminal to output the sound of the display device 200.
[0067] In some embodiments, the user input interface 280 can be used to receive instructions from the user input.
[0068] To perform user interactions, in some embodiments, the display device 200 can run an operating system. The operating system is a computer program for managing and controlling hardware resources and software resources in the display device 200. The operating system can control the display device 200 to provide a user interface, for example, the operating system can directly control the display device 200 to provide a user interface, or can provide a user interface by running an application program. The operating system also allows the user to interact with the display device 200.
[0069] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for the display device 200.
[0070] The operating system can be divided into different modules or levels according to the implemented functions, for example, as shown in Figure 3 In some embodiments, the system is divided into four layers, from top to bottom, the application layer (referred to as the "application layer"), the application framework layer (referred to as the "framework layer"), the system library layer, and the kernel layer.
[0071] In some embodiments, the application layer is used to provide services and interfaces for applications, so that the display device 200 can run the application and interact with the user based on the application. At least one application can run in the application layer, which can be a window (Window) program, a system setting program or a clock program provided by the operating system, or an application developed by a third-party developer. In specific implementation, the application package in the application layer is not limited to the above examples.
[0072] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications. The application framework layer includes some pre-defined functions. The application framework layer is equivalent to a processing center that decides which application in the application layer to act. The application can access the resources in the system and obtain the services of the system through the API interface during execution.
[0073] As shown in Figure 3As shown, the application framework layer in this embodiment includes a view system, managers, and content providers. The view system designs and implements the application's interface and interactions, and includes lists, grids, text boxes, and buttons. The managers include at least one of the following modules: an activity manager for interacting with all running activities in the system; a location manager for providing system services or applications with access to system location services; a package manager for retrieving various information related to application packages currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.
[0074] In some embodiments, the Activity Manager manages the lifecycle of individual applications and common navigation and back functions, such as controlling application exit, opening, and back actions. The Window Manager manages all window programs, such as obtaining the screen size, determining if a status bar is present, locking the screen, capturing the screen, and controlling changes to the display window, such as shrinking the display window, shaking the display, or distorting the display.
[0075] In some embodiments, the system runtime library layer can provide support for the framework layer. When the framework layer is used, the operating system runs the instruction library contained in the system runtime library layer, such as the C / C++ instruction library, to implement the functions to be performed by the framework layer.
[0076] In some embodiments, the kernel layer is a functional layer situated between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. For example, ... Figure 3 As shown, hardware drivers can be configured in the kernel layer. The drivers included in the kernel layer can be at least one of the following: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.
[0077] It should be noted that the above examples are only simple divisions of the operating system functions, and do not constitute limitations on the specific operating system forms of the display device 200 in the embodiments of the present application. Depending on the functions of the display device, the type of the operating system, and other factors, the number and specific types of the levels included in the operating system can take other forms.
[0078] In some embodiments, the display device 200 can display a specific user interface in response to a specific control instruction input by the user. For example, the display device 200 can display a home interface in response to a power-on instruction input by the user.
[0079] For another example, the display device 200 can display a media content playing interface in response to a media content playing instruction input by the user. In some embodiments, the media content played by the display device 200 through the media content playing interface can be media content from different channels, including media content in a video platform, such as pre-stored media content in "Video Platform A", live video in a live broadcast software, such as live video in "Sports Channel", and user-made media content, such as a video composed of two original videos by the user.
[0080] In some embodiments, the display device 200 can control a media content composition application installed in the display device 200 to compose media content data in response to a media content composition instruction input by the user. As shown in Figure 4 When composing media content data, the display device 200 can first perform step S401: obtaining original media content data to be composed. For example, the controller 250 in the display device 200 can shoot a video through a built-in camera to obtain original media content data. For another example, the controller 250 can also receive a video uploaded by the user as original media content data.
[0081] After obtaining the original media content data, the controller 250 performs step S402: inputting the original media content data to be composed into the media content composition application. For example, the display device 200 obtains first original media content data and second original media content data, and then inputs the first original media content data and the second original media content data into the media content composition application for processing.
[0082] Finally, the controller 250 performs step S403: controlling the media content composition application to compose the original media content data to be composed into new media content data. For example, the display device 200 inputs the first original media content data and the second original media content data into the media content composition application, and controls the media content composition application to perform processing, including cropping, splicing, and composition, on the first original media content data and the second original media content data to generate new media content data.
[0083] However, the raw media asset data that the display device 200 can obtain is limited, making it difficult to accurately acquire the corresponding raw media asset data according to user needs, resulting in low richness of the synthesized media asset data. Furthermore, the process of synthesizing media asset data through media asset compositing applications by the display device 200 first requires acquiring the raw media asset data and then processing it, making the synthesis speed slow.
[0084] In some embodiments, the display device 200 may also execute a media asset compositing method to composite media assets, thereby addressing the problems of slow speed and low richness in media asset compositing by the display device 200. To facilitate the implementation of the media asset compositing method, the display device 200 should include at least a display 260 and a controller 250. The display 260 is configured to display a user interface; such as... Figure 5 As shown, the controller 250 is configured to execute the program steps corresponding to the media asset synthesis method, including the following:
[0085] S100: Responds to the user's input media asset compositing command and obtains the original media asset data.
[0086] In some embodiments, such as Figure 6 As shown, after receiving the media asset compositing instruction input by the user, the controller 250 first acquires the raw media asset data. This raw media asset data includes data in various formats, such as raw image data, raw audio data, and raw text data.
[0087] In some embodiments, the controller 250 may acquire a raw image data, a raw audio data, and a raw text data as the same set of raw media asset data for generating a composite media asset data. For example, in response to a media asset compositing instruction input by a user, the controller 250 acquires a "first image," a "first audio," and a "first text," and uses the "first image," "first audio," and "first text" to generate composite media asset data.
[0088] In some embodiments, the playback duration of the synthesized media asset data is equal to the playback duration of the original audio data. Therefore, by obtaining the playback duration of the original audio data, the playback duration of the synthesized media asset data can be obtained. When the controller 250 obtains the playback duration of the original audio, it can first obtain the data length, sampling frequency, number of channels, and bit depth of the original audio data. For example, the controller 250 obtains the data length of the "first audio" as 1,048,576 bytes; the sampling frequency as 44,100 Hz; the number of channels as 2; and the bit depth as 16 bits.
[0089] Then, calculate the playback duration of the original audio data based on the data length, sampling frequency, number of channels, and bit depth. For example, according to the formula: The total number of sample points for the "first audio" is calculated as follows: One; then according to the formula: The playing duration of the "first audio" is calculated as
[0090] In some embodiments, as shown in FIG. 2B, the media content synthesis architecture of the controller 250 includes a data buffer module. If the raw media content data continuously acquired by the display device 200 exceeds the maximum data amount that the controller 250 can process in real time, the raw media content data that cannot be processed in real time can be stored in the data buffer module. Figure 7
[0091] For example, the maximum data amount that the controller 250 can process in real time is 60 megabytes (MB), and the raw media content data acquired by the controller 250 in response to the media content synthesis instruction input by the user is "first picture", "first audio", "first text", "second picture", and "second audio". The data amount of the "first picture" is 30 MB, the data amount of the "first audio" is 20 MB, the data amount of the "first text" is 10 MB, the data amount of the "second picture" is 40 MB, and the data amount of the "second audio" is 10 MB. The total data amount of all the raw media content data acquired by the controller 250 is greater than the maximum data amount that the controller 250 can process in real time. At this time, the controller 250 processes the "first picture", "first audio", and "first text" in the order in which the raw media content data is acquired, and stores the "second picture" and "second audio" that cannot be processed in real time in the data buffer module in the order in which the raw media content data is acquired, and then processes the "second picture" and "second audio" in the order in which the raw media content data is stored.
[0092] S200: performing region-wise extraction on the raw picture data to obtain video frame data.
[0093] In some embodiments, as shown in FIG. 2B, the media content synthesis architecture of the controller 250 includes a data buffer module. If the raw media content data continuously acquired by the display device 200 exceeds the maximum data amount that the controller 250 can process in real time, the raw media content data that cannot be processed in real time can be stored in the data buffer module. Figure 7 In some embodiments, if the data format of the decoded picture data is different from the preset format, data format conversion is performed on the decoded picture data to obtain decoded picture data in the preset format. For example, the preset format is set as YUV format, and if the data format of the decoded picture data obtained by the controller 250 after decoding the "first picture" is RGB format, the data format of the decoded picture data needs to be converted to YUV format.
[0094]
[0095] In some embodiments, after acquiring decoded image data in a preset format, the controller 250 can perform region-based extraction on the decoded image data to obtain video frame data. For example... Figure 8 As shown, when the controller 250 performs region extraction on the decoded image data, it can first execute step S801: obtaining the first boundary point and the second boundary point of the decoded image data. The first boundary point and the second boundary point are two symmetrical vertices in the decoded image data. For example, Figure 9 This is a schematic diagram of the "first image" obtained by controller 250. Figure 9 It can be seen that the width of the "first image" is 1920 pixels and the height is 1680 pixels. A can be used as the first boundary point and B as the second boundary point.
[0096] Next, the controller 250 executes step S802: setting an extraction window based on the first boundary point and the second boundary point. The width of the extraction window is equal to a preset ratio of the number of pixels contained between the first and second boundary points. In some embodiments, the extraction window is used to extract data from the decoded image data by region, thereby splitting the decoded image data into multiple video frame data, and then stitching the multiple video frame data together into video data. For example, if the preset ratio is set to 80%, and the width of the "first image" is 1920 pixels, then the width of the extraction window is 1920 * 80% = 1536 pixels.
[0097] In some embodiments, to improve the completeness of video frame data extraction, the extraction window can be set to a rectangle, and the height of the extraction window can be the same as the height of the decoded image data. For example, if the height of the "first image" is 1680 pixels, then the height of the extraction window is set to 1680 pixels. The extraction window finally determined by the controller 250 is as follows: Figure 10 As shown in the dashed box in the image.
[0098] It should be noted that the extraction window can be set to various shapes, such as rectangles, squares, etc., as long as it can completely extract video frames from the decoded image data. No specific limitation is made in this embodiment.
[0099] In some embodiments, after the controller 250 has determined the shape and size of the first boundary point, the second boundary point, and the extraction window, it can begin executing step S803: using the first boundary point as the starting point of the extraction window, and using a preset number as the translation step size of the extraction window, extracting pixels located within the extraction window sequentially in the decoded image data in a first direction to obtain multiple video frame data. For example, as Figure 10 As shown, the first boundary point A is taken as the starting point of the extraction window, the preset number is set to 1 pixel, and the first direction is set to the direction from the first boundary point A to the second boundary point B, that is... Figure 10The extraction window extracts the first video frame data of the first stage, and then moves one pixel point in the direction of the second boundary point B. The extraction window extracts the second video frame data of the first stage, and so on.
[0100] In some embodiments, each video frame data includes an extraction boundary point, i.e., a boundary point determined by the extraction window. For example, as shown in FIG. 3, if the first boundary point of the decoded picture data is A and the second boundary point is B, the extraction boundary point of the video frame data in the first stage is the point closest to the second boundary point B. Figure 10
[0101] In some embodiments, during the extraction of the video frame data, the extraction window gradually moves towards the second boundary point by a preset number of pixel points as a translation step. If the extraction boundary point of the video frame data coincides with the second boundary point, the second stage is entered. In the second stage, the second boundary point is taken as the starting point of the extraction window, the preset number is taken as the translation step of the extraction window, and the pixel points in the extraction window are extracted in the decoded picture data in the second direction one by one to obtain a plurality of video frame data. For example, as shown in FIG. 4, the second boundary point B is taken as the starting point of the extraction window, the preset number is set to one pixel point, and the second direction is set to the direction from the second boundary point B to the first boundary point A, i.e. Figure 11 Figure 11 The extraction window extracts the first video frame data of the first stage, and then moves one pixel point in the direction of the second boundary point B. The extraction window extracts the second video frame data of the first stage, and so on.
[0102] It should be noted that, since the moving direction of the extraction window in the second stage is opposite to that in the first stage, the extraction boundary point of the video frame data in the second stage is also opposite to that in the first stage, as shown in FIG. 5, which is the point closest to the first boundary point A. Figure 11
[0103] In some embodiments, in the second stage, when the extraction window moves to the point where the extraction boundary point of the video frame data coincides with the first boundary point A, it is indicated that the second stage of the video frame data extraction is completed, i.e., the controller 250 has extracted the required plurality of video frame data.
[0104] S300: encoding the video frame data according to a preset encoding format to obtain video data.
[0105] In some embodiments, after extracting video frame data, the controller 250 can calculate the time interval between video frame data according to the set frame rate and send it to the encoder. The encoder then encodes the video frame data to obtain the video data required for synthesizing media asset data. For example, if the frame rate set by the controller 250 is 30 frames per second, the time interval between two video frame data is 0.1332 seconds. The video frame data and the time interval are sent to the encoder, which encodes the video frame data into H.264 format. The H.264 format can provide high-quality image transmission under low bandwidth and has a high compression ratio, which can improve image quality while reducing the resources occupied by image storage and transmission.
[0106] It should be understood that video frame data can also be encoded in other formats besides H.264, as long as it meets the requirements for synthesizing media asset data. No specific limitations are made in this application embodiment.
[0107] In some embodiments, the controller 250 may sequentially process all the original image data in the data cache module. For example... Figure 12 As shown, when controller 250 generates video data based on raw image data from the data cache module, it first decodes the raw image data and then determines whether the data format of the decoded image data is a preset format. If the data format of the decoded image data is not a preset format, it is converted to the preset format. Next, it performs the first stage of region extraction on the decoded image data in the preset format. During region extraction, after extracting each video frame, controller 250 encodes the video frame data and then moves the extraction window along the first direction to extract the next video frame data, while simultaneously determining whether the second boundary point coincides with the extraction boundary point. If the second boundary point coincides with the extraction boundary point, the second stage of region extraction is performed, moving the extraction window along the second direction to extract the next video frame data. During the second stage of region extraction, controller 250 determines whether the first boundary point coincides with the extraction boundary point. If the first boundary point coincides with the extraction boundary point, controller 250 has initially completed the processing of one raw image data. Finally, controller 250 determines whether region extraction is complete based on the video frame rate of the region extraction.
[0108] In some embodiments, after the controller 250 finishes processing one piece of original image data in the data cache module, it determines whether there is still unprocessed original image data in the data cache module. If so, it continues to process the original image data according to the above method.
[0109] As can be seen from the above embodiments, performing sub-region extraction from the first boundary point A to the second boundary point B on the decoded picture data can generate video data with left-right moving effect. It should be understood that if the positions of the first boundary point and the second boundary point are changed, video data with different moving effects can be generated. As shown in Figure 9 , if A is set as the first boundary point and C is set as the second boundary point, video data with up-down moving effect can be generated, which is not specifically limited in the embodiments of the present application.
[0110] S400: Perform format conversion processing on the original audio data to obtain audio data in a preset encoding format.
[0111] In some embodiments, as shown in Figure 7 , the media synthesis architecture of the controller 250 further includes an audio conversion module. When the controller 250 obtains the original audio data in response to the media synthesis instruction input by the user, the original audio data can be input to the audio conversion module to perform decoding, and then the decoded audio data is encoded to obtain audio data in a preset encoding format.
[0112] For example, the original audio data obtained by the controller 250 is "first audio". First, the "first audio" is decoded. The decoder decompresses the compressed "first audio" into original audio sampling points according to the encoding format and compression algorithm of the "first audio". Then, the audio sampling points are quantized, and the values of the audio sampling points are mapped to a discrete set of values, and binary digital signal values are used to represent. Finally, the quantized digital signal values are converted into a preset format suitable for storage or transmission, such as AAC format, to obtain audio data in a preset encoding format.
[0113] It should be understood that the audio data can also be encoded into other formats in addition to AAC format, as long as it meets the requirements for synthesizing media data, which is not specifically limited in the embodiments of the present application.
[0114] In some embodiments, the controller 250 can also obtain new audio data and combine the new audio data with the original audio data to obtain synthesized audio data, so as to improve the richness of the synthesized media data.
[0115] As shown in Figure 13 , when generating the synthesized audio data, the controller 250 can further perform step S1301 of obtaining new audio data. For example, when the controller 250 synthesizes media data using "first audio", it also obtains "second audio" as new audio data.
[0116] After acquiring the original audio data and the newly added audio data, controller 250 executes step S1302: decoding the original audio data to obtain first decoded audio data. And executes step S1303: decoding the newly added audio data to obtain second decoded audio data. For example, controller 250 acquires "first audio" as the original audio data and acquires "second audio" as the newly added audio data. It decodes the "first audio" to obtain first decoded audio data, and decodes the "second audio" to obtain second decoded audio data.
[0117] After acquiring the first decoded audio data and the second decoded audio data, the controller 250 executes step S1304: combining the first decoded audio data and the second decoded audio data to obtain synthesized audio data. For example, the controller 250 combines the first decoded audio data obtained by decoding the "first audio" and the second decoded audio data obtained by decoding the "second audio" into unified synthesized audio data.
[0118] Finally, the controller 250 executes step S1305: encoding the synthesized audio data to obtain audio data in a preset encoding format. For example, if the preset format is AAC, the controller 250 can obtain AAC format audio data by encoding the synthesized audio data through an encoder. The audio data can also be encoded in other formats besides AAC, but this embodiment does not impose specific limitations.
[0119] S500: Adds time stamps to the raw text data to obtain caption data.
[0120] In some embodiments, such as Figure 7 As shown, the media asset compositing architecture of controller 250 also includes a text conversion module. When controller 250 responds to a user-inputted media asset compositing command and obtains the original text data, it can input the original text data into the text conversion module. The text conversion module adds time stamps to the original text data to obtain subtitle data.
[0121] In some embodiments, such as Figure 14 As shown, when the controller 250 adds time stamps to the original text data through the text conversion module, it can first execute step S1401: obtain the sampling frequency of the original audio data in the same group as the original text data and the number of sampling points in each frame of audio data. For example, if the original text data obtained by the controller 250 is "first text" and the original audio data in the same group as "first text" is "first audio", then the controller 250 obtains the sampling frequency of "first audio" as 44100 Hz; and the number of sampling points in each frame of audio data as 1024.
[0122] Secondly, the controller 250 performs step S1402: calculates the starting time point and the ending time point of the original audio data according to the sampling frequency and the sampling point number of each frame of audio data. For example, the sampling frequency of the "first audio" obtained by the controller 250 is 44100 Hz, and the sampling point number of each frame of audio data is 1024, and then the time interval of each frame of audio data in the "first audio" is calculated according to the formula: The time interval of each frame of audio data in the "first audio" is calculated as 0.0232 seconds, and then the starting time point and the ending time point of each frame of audio data are determined, such as the starting time point of the first frame of audio data is 0 seconds, and the ending time point is 0.0232 seconds, the starting time point of the second frame of audio data is 0.0232 seconds, and the ending time point is 0.0464 seconds, and so on.
[0123] Finally, the controller 250 performs step S1403: adds the starting time point and the ending time point of the original audio data to the original text data to obtain the subtitle data. For example, the original audio data obtained by the controller 250 is "first audio", and the original text data obtained is "first text", and the starting time point and the ending time point of each piece of text content in the "first text" are added according to the starting time point and the ending time point of each frame of audio data in the "first audio", such as the "first audio" includes "first audio frame", "second audio frame" and "third audio frame", and if the time interval of each frame of audio data in the "first audio" is 2.6 seconds, the "first text" includes "first sentence", "second sentence" and "third sentence", then the "first sentence" is added with the starting time point of 0 seconds and the ending time point of 2.6 seconds, the "second sentence" is added with the starting time point of 2.6 seconds and the ending time point of 5 seconds, and the "third sentence" is added with the starting time point of 5 seconds and the ending time point of 7.5 seconds.
[0124] In some embodiments, when the starting time point and the ending time point are added to the original text data, the subtitle data in the preset format can be generated. For example, after the controller 250 adds the starting time point and the ending time point to the "first sentence", "second sentence" and "third sentence" in the "first text", firstly, the style of the subtitle data is defined, including font, size, color, border and the like. Secondly, the format of the subtitle data is set as ASS format. Thirdly, the playing resolution suitable for the video data is defined. Finally, the subtitle style is edited to generate the subtitle data meeting the user's demand. It should be understood that the subtitle data can also be in other formats except for the ASS format, as long as it meets the requirement for synthesizing the media data, which is not limited in the embodiments of the present application.
[0125] In some embodiments, each subtitle data generated by the controller 250 cannot be fully displayed in the corresponding video frame data. Therefore, to improve the consistency between the subtitle data and the video data, and to ensure that the subtitle data can be displayed correctly when the user controls the controller 250 to jump to a specified video screen, the controller 250 needs to perform frame interpolation processing on the subtitle data. First, the start time marker of the subtitle data is obtained. For example, the subtitle data generated by the controller 250 from "first text" includes subtitle data corresponding to "first statement", subtitle data corresponding to "second statement", and subtitle data corresponding to "third statement". Taking the subtitle data corresponding to "first statement" as an example, the time marker is a start time point of 0 seconds and an end time point of 2.6 seconds.
[0126] Starting from the time indicated by the start time marker, subtitle data is added to the initial media asset template at preset time intervals. For example, if the start time of the subtitle data corresponding to "First Statement" is 0 seconds, and the preset time interval is 500 milliseconds, then starting from 0 seconds, subtitle data corresponding to "First Statement" is inserted every 0.5 seconds between 0 seconds and 2.6 seconds, for a total of 5 subtitle data corresponding to "First Statement": "First Inserted Subtitle," "Second Inserted Subtitle," "Third Inserted Subtitle," "Fourth Inserted Subtitle," and "Fifth Inserted Subtitle." The start time of "First Inserted Subtitle" is 0.5 seconds, "Second Inserted Subtitle" is 1 second, "Third Inserted Subtitle" is 1.5 seconds, "Fourth Inserted Subtitle" is 2 seconds, and "Fifth Inserted Subtitle" is 2.5 seconds.
[0127] S600: Performs encapsulation processing on video data, audio data, and subtitle data to obtain composite media asset data.
[0128] In some embodiments, such as Figure 7 As shown, the media asset compositing architecture of controller 250 also includes a data encapsulation module. After controller 250 converts the original image data, original audio data, and original text data into video data, audio data in a preset encoding format, and subtitle data, the data encapsulation module can perform encapsulation processing to obtain the composite media asset data. For example, controller 250 converts the "first image" into H.264 format video data, the "first audio" into AAC format audio data, and the "first text" into ASS format subtitle data, and inputs these into the data encapsulation module for encapsulation processing to generate composite media asset data.
[0129] In some embodiments, such as Figure 15As shown, when the controller 250 generates the synthesized media data, it can first perform step S1501: generating an initial media template. For example, the controller 250 needs to generate the synthesized media data in the MP4 format in response to the media synthesis instruction input by the user. Then in the data packaging module, the initial media template is first generated, and the preset media format is set to the MP4 format in the initial media template. It should be understood that the preset media format can also be other formats in addition to the MP4 format, which is not limited in the embodiments of the present application.
[0130] Secondly, the controller 250 performs step S1502: inputting the video data, the audio data and the subtitle data into the initial media template. In some embodiments, as shown in FIG. 15B, the controller 250 can perform step S1502 in the following way. Figure 7 As shown, the media synthesis architecture of the controller 250 further includes a video transmission pipeline, an audio transmission pipeline and a subtitle transmission pipeline. Among them, the video transmission pipeline, the audio transmission pipeline and the subtitle transmission pipeline correspond to one data transmission thread respectively. When the controller 250 converts the original media data into the video data, the audio data and the subtitle data through the picture conversion module, the audio conversion module and the text conversion module, the video data can be transmitted through the data transmission thread corresponding to the video transmission pipeline, the audio data can be transmitted through the data transmission thread corresponding to the audio transmission pipeline, and the subtitle data can be transmitted through the data transmission thread corresponding to the subtitle transmission pipeline, and then the converted video data, the audio data and the subtitle data are input into the data packaging module in real time. It should be understood that by respectively setting different transmission pipelines to transmit the video data, the audio data and the subtitle data, the transmission speed of the data can be improved, and then the speed of the controller 250 generating the synthesized media data can be improved.
[0131] Finally, the controller 250 performs step S1503 and step S1504: obtaining an end identifier, and based on the end identifier, generating the synthesized media data in the preset media format according to the initial media template. In some embodiments, the controller 250 can perform data conversion on multiple groups of original media data, and encapsulate the converted data as one synthesized media data. When the controller 250 obtains the end identifier in the data encapsulation module, it indicates that the encapsulation process is completed, and the controller 250 no longer obtains new original media data. For example, the controller 250 obtains two groups of original media data, the first group of original media data including "first picture", "first audio" and "first text", and the second group of original media data including "second picture", "second audio" and "second text". According to the order of obtaining, the conversion is performed through the picture conversion module, the audio conversion module and the text conversion module to obtain two groups of converted video data, audio data and subtitle data. When the controller 250 sends the two groups of converted video data, audio data and subtitle data to the data encapsulation module through the video transmission pipeline, the audio transmission pipeline and the subtitle transmission pipeline, the data encapsulation module detects the end identifier, and then the data encapsulation module encapsulates the two groups of video data, audio data and subtitle data as one synthesized media data according to the input order.
[0132] It should be noted that in the media synthesis architecture of the controller 250, the process of converting the original media data by the picture conversion module, the audio conversion module and the text conversion module and the process of encapsulating the converted data by the data encapsulation module can be performed asynchronously, that is, when the picture conversion module, the audio conversion module and the text conversion module perform conversion on the second group of original media data, the data encapsulation module can perform encapsulation on the converted data of the first group of original media data. Therefore, when the controller 250 processes multiple groups of original media data using the media synthesis architecture, the speed of generating synthesized media data can be improved.
[0133] In some embodiments, the original media data obtained by the controller 250 in response to the media synthesis instruction input by the user further includes original video data. When the controller 250 obtains the original video data, it first parses the original video data into original picture data, original audio data and original text data. Then, the conversion is performed through the picture conversion module, the audio conversion module and the text conversion module to obtain video data, audio data and subtitle data. Finally, the video data, the audio data and the subtitle data are transmitted to the data encapsulation module through the video transmission pipeline, the audio transmission pipeline and the subtitle transmission pipeline for encapsulation to obtain the synthesized media data.
[0134] Based on the display device 200, the application further provides a media synthesis method, comprising: in response to a media synthesis instruction input by a user, obtaining original media data, wherein the original media data comprises original picture data, original audio data and original text data;
[0135] performing regional extraction on the original picture data in units of pixels to obtain video frame data;
[0136] performing encoding on the video frame data according to a preset encoding format to obtain video data;
[0137] performing format conversion processing on the original audio data to obtain audio data in the preset encoding format;
[0138] adding time labels to the original text data to obtain subtitle data;
[0139] performing packaging processing on the video data, the audio data and the subtitle data to obtain synthesized media data; wherein a playing time length of the synthesized media data is equal to a playing time length of the original audio data.
[0140] According to the above technical solution, the application provides a display device and a media synthesis method. The method can obtain original media data in response to a media synthesis instruction input by a user; perform regional extraction on original picture data in units of pixels to obtain video frame data; perform encoding on the video frame data according to a preset encoding format to obtain video data; perform format conversion processing on the original audio data to obtain audio data in the preset encoding format; add time labels to the original text data to obtain subtitle data; and perform packaging processing on the video data, the audio data and the subtitle data to obtain synthesized media data. The method can obtain various types of original media data, perform corresponding processing on the obtained original media data in real time, and then perform packaging on the processed media data to obtain synthesized media data, thereby improving the speed and richness of the synthesized media data.
[0141] The same or similar parts among the various embodiments in the specification can be referred to each other and will not be described here.
[0142] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present application can be implemented by means of software plus necessary universal hardware platforms. Based on such understanding, the technical solutions in the embodiments of the present application can be embodied in the form of a software product, and the computer software product can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a plurality of instructions to cause a computer device (which can be a personal computer, a server, or a network device, and the like) to execute the methods of the various embodiments or some parts of the embodiments of the present application.
[0143] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
[0144] For the convenience of explanation, the above description has been made in combination with specific embodiments. However, the above exemplary discussion is not intended to exhaust or limit the embodiments to the specific forms disclosed above. Various modifications and variations can be derived according to the above teachings. The selection and description of the above embodiments are to better explain the principles and practical applications, so that those skilled in the art can better use the embodiments and various different modified embodiments suitable for specific use considerations.
Claims
1. A display device, characterized by comprising: The method comprises the following steps: a display configured to display a user interface; a controller configured to: in response to a user inputted media synthesis instruction, acquire original media data, the original media data comprising original picture data, original audio data and original text data; perform decoding on the original picture data to obtain decoded picture data; acquire a first boundary point and a second boundary point of the decoded picture data, the first boundary point and the second boundary point being two symmetrical vertices in the decoded picture data; set an extraction window according to the first boundary point and the second boundary point, the width of the extraction window being equal to a preset proportion of the pixel points contained between the first boundary point and the second boundary point; extract pixel points in the decoded picture data in the extraction window in a first direction with the first boundary point as the starting point of the extraction window and a preset number as the translation step of the extraction window to obtain a plurality of video frame data, the first direction being from the first boundary point to the second boundary point; the video frame data comprising an extraction boundary point; in response to the extraction boundary point of the video frame data coinciding with the second boundary point, extract pixel points in the decoded picture data in the extraction window in a second direction with the second boundary point as the starting point of the extraction window and the preset number as the translation step of the extraction window to obtain a plurality of the video frame data, the second direction being from the second boundary point to the first boundary point; perform encoding on the video frame data in a preset encoding format to obtain video data; perform format conversion processing on the original audio data to obtain audio data in a preset encoding format; add time markers to the original text data to obtain subtitle data; perform packaging processing on the video data, the audio data and the subtitle data to obtain synthesized media data; wherein the playing time length of the synthesized media data is equal to the playing time length of the original audio data.
2. The display device of claim 1, wherein, Before the controller performs the step of performing first-stage regional extraction on the decoded picture data in units of pixels, the controller is further configured to: if the data format of the decoded picture data is different from a preset format, perform data format conversion on the decoded picture data to obtain decoded picture data in a preset format.
3. The display device of claim 1, wherein, The controller performs the step of performing format conversion processing on the original audio data to obtain audio data in a preset encoding format, specifically configured to: acquire new audio data; perform decoding on the original audio data to obtain first decoded audio data; perform decoding on the new audio data to obtain second decoded audio data; synthesize the first decoded audio data and the second decoded audio data to obtain synthesized audio data; perform encoding on the synthesized audio data to obtain audio data in a preset encoding format.
4. The display device of claim 1, wherein, The controller performs the step of adding time markers to the original text data to obtain subtitle data, specifically configured to: acquire the sampling frequency and the number of sampling points of the original audio data; calculating a start time point and an end time point of the original audio data according to the sampling frequency and the number of sampling points; adding the start time point and the end time point of the original audio data to the original text data to obtain the subtitle data.
5. The display device of claim 4, wherein, The controller is configured to perform encapsulation processing on the video data, the audio data and the subtitle data to obtain the synthesized media data, and specifically configured to: generate an initial media template, the initial media template including a preset media format; input the video data, the audio data and the subtitle data into the initial media template; obtain an end identifier; generate the synthesized media data in the preset media format according to the initial media template based on the end identifier.
6. The display device of claim 5, wherein, After the step of inputting the video data, the audio data and the subtitle data into the initial media template, the controller is further configured to: obtain a start time mark of the subtitle data; add the subtitle data in the initial media template every preset time interval with the time point represented by the start time mark as a starting point.
7. The display device of claim 1, wherein, After the step of obtaining the original media data in response to the media synthesis instruction input by the user, the controller is further configured to: obtain a data length, a sampling frequency, a number of channels and a bit depth of the original audio data; calculate a playing duration of the original audio data according to the data length, the sampling frequency, the number of channels and the bit depth.
8. A method of media composition, the method comprising: The method is applied to the display device of any one of claims 1-7, and the method comprises: obtaining original media data in response to a media synthesis instruction input by a user, the original media data including original picture data, original audio data and original text data; performing decoding on the original picture data to obtain decoded picture data; obtaining a first boundary point and a second boundary point of the decoded picture data, the first boundary point and the second boundary point being two vertices symmetrically located in the decoded picture data; setting an extraction window according to the first boundary point and the second boundary point, a width of the extraction window being equal to a preset proportion of pixel points contained between the first boundary point and the second boundary point; extracting pixel points located in the extraction window in the decoded picture data in a first direction with the first boundary point as a starting point of the extraction window and a preset number as a translation step of the extraction window to obtain a plurality of video frame data, the first direction being from the first boundary point to the second boundary point; the video frame data including an extraction boundary point; extracting pixel points located in the extraction window in the decoded picture data in a second direction with the second boundary point as a starting point of the extraction window and the preset number as a translation step of the extraction window in response to the extraction boundary point of the video frame data coinciding with the second boundary point, the second direction being from the second boundary point to the first boundary point; performing encoding on the video frame data in a preset encoding format to obtain video data; performing format conversion processing on the original audio data to obtain audio data in a preset encoding format; and performing format conversion processing on the original audio data to obtain audio data in a preset encoding format. adding time marks to the original text data to obtain subtitle data; performing packaging processing on the video data, the audio data and the subtitle data to obtain synthesized media asset data; wherein a playing time length of the synthesized media asset data is equal to a playing time length of the original audio data.
Citation Information
Patent Citations
Multi-material video synthesis method and device, electronic equipment and storage medium
CN112153463A