File playing method and device, equipment, storage medium and program product

By using virtual input components in cloud mobile phones to play the generated files to be played for the target device, the problems of high cost of cloud mobile phones and poor user experience are solved, and lower cost and convenient file provision and content release are achieved.

CN120050356APending Publication Date: 2025-05-27GUANGZHOU DULING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510122549.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

How to reduce the cost of users using cloud mobile phones and improve users' experience of cloud mobile phones, especially in file provision and content release.

Method used

By generating a file to be played in response to the received file description information, and establishing a communication connection with the device corresponding to the communication initiation request, the file to be played for the target device is played by the virtual input component, and the virtual input device is set in the cloud mobile phone through virtualization technology.

Benefits of technology

It reduces the file provision cost when users use cloud mobile phones, and reduces the communication requirements between the terminal devices actually used by users and the cloud mobile phones, so that users can use cloud mobile phones to provide and publish content at a lower cost and more convenient manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050356A_ABST
    Figure CN120050356A_ABST
Patent Text Reader

Abstract

The invention provides a file playing method and device, electronic equipment, a computer readable storage medium and a computer program product, is applied to a cloud mobile phone, and relates to the technical field of artificial intelligence such as cloud data, online playing and content generation. A specific embodiment of the method comprises the steps of generating a corresponding to-be-played file based on file description information in response to the received file description information sent by a first device; in response to the received communication initiation request of the first device, establishing a communication connection with a second device corresponding to the communication initiation request; the to-be-played file is played for the second device through the virtual input component, and the virtual input device is arranged in the cloud mobile phone through the virtualization technology. Therefore, the file providing cost when the user uses the cloud mobile phone can be reduced, and the communication requirement between the terminal equipment actually used by the user and the cloud mobile phone can be reduced, so that the user can use the cloud mobile phone to provide and publish the content more conveniently at lower cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, specifically to cloud phones, and involves artificial intelligence technology fields such as cloud data, online playback, and content generation. In particular, it relates to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for playing files. Background Art

[0002] With the development of computer technologies, in order to improve the computing power and user experience on the user device side, cloud phone technology has emerged. A cloud phone is to apply cloud computing technology to network terminal services and realize a "phone" in the form of cloud services through a cloud server. For example, for a user, they can use their terminal device as an operation carrier and a "server" as a device providing cloud computing power, and interact with the "server" through the terminal device to realize a "cloud phone".

[0003] This "phone" that deeply combines network services can rely on its own system and the network terminals set up by manufacturers to obtain more powerful computing power and provide more functions for the user's phone. Therefore, how to reduce the usage cost of cloud phones for users and improve the user experience of cloud phones is worthy of attention and an urgent need. Summary of the Invention

[0004] Embodiments of the present disclosure propose a method, apparatus, electronic device, computer-readable storage medium, and computer program product for playing files.

[0005] In a first aspect, embodiments of the present disclosure propose a method for playing a file, including: in response to receiving file description information sent by a first device, generating a corresponding file to be played based on the file description information; in response to receiving a communication initiation request from the first device, establishing a communication connection with a second device corresponding to the communication initiation request; playing the file to be played for the second device through a virtual input component, where the virtual input device is set in a cloud phone through virtualization technology.

[0006] In a second aspect, embodiments of the present disclosure propose an apparatus for playing a file, including: a to-be-played file generation unit configured to, in response to receiving file description information sent by a first device, generate a corresponding file to be played based on the file description information; a first communication connection establishment unit configured to, in response to receiving a communication initiation request from the first device, establish a communication connection with a second device corresponding to the communication initiation request; a to-be-played file playback unit configured to play the file to be played for the second device through a virtual input component, where the virtual input device is set in a cloud phone through virtualization technology.

[0007] In a third aspect, embodiments of the present disclosure provide an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to implement the method for playing a file described in any implementation manner of the first aspect.

[0008] In a fourth aspect, embodiments of the present disclosure provide a non-transitory computer-readable storage medium storing computer instructions, and when the computer instructions are executed by a computer, the computer is enabled to implement the method for playing a file described in any implementation manner of the first aspect.

[0009] In a fifth aspect, embodiments of the present disclosure provide a computer program product including a computer program, and when the computer program is executed by a processor, the computer program is enabled to implement the method for playing a file described in any implementation manner of the first aspect.

[0010] The method, device, electronic device, computer-readable storage medium, and computer program product for playing a file provided by embodiments of the present disclosure, in response to receiving file description information sent by a first device, generate a corresponding file to be played based on the file description information; in response to receiving a communication initiation request from the first device, establish a communication connection with a second device corresponding to the communication initiation request; and play the file to be played for the second device through a virtual input component, wherein the virtual input device is set in a cloud phone through virtualization technology.

[0011] The present disclosure can not only reduce the file providing cost when a user uses a cloud phone, but also reduce the communication requirements between the terminal device actually used by the user and the cloud phone, enabling the user to more conveniently provide and publish content using the cloud phone at a lower cost.

[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present disclosure will become more apparent:

[0014] Figure 1 is an exemplary system architecture to which the present disclosure can be applied;

[0015] Figure 2 is a flowchart of a method for playing a file provided by an embodiment of the present disclosure;

[0016] Figure 3A flowchart of a process for generating a file to be played provided by an embodiment of the present disclosure;

[0017] Figure 4 A flowchart of a process for playing a file in an application scenario provided by an embodiment of the present disclosure;

[0018] Figure 5 A structural block diagram of a device for playing a file provided by an embodiment of the present disclosure;

[0019] Figure 6 A schematic structural diagram of an electronic device suitable for executing a method for playing a file provided by an embodiment of the present disclosure. Detailed implementation manners

[0020] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below. It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0021] In addition, in the technical solutions involved in the present disclosure, the acquisition, storage, use, processing, transportation, provision, and disclosure, etc., of the information involved (for example, user personal information) all comply with the provisions of relevant laws and regulations and do not violate public order and good customs. For example, the acquisition and use of the file description information provided by the user are known, authorized, and permitted by the user.

[0022] Figure 1 An exemplary system architecture 100 is shown in which the method, device, electronic device, and computer-readable storage medium for playing a file according to the present disclosure can be applied.

[0023] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium to provide a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0024] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications for implementing information communication between the two can be installed on terminal devices 101, 102, 103 and server 105, such as cloud phone call applications, cloud service-based file playback applications, instant messaging applications, etc.

[0025] Terminal devices 101, 102, 103 and server 105 can be either hardware or software. When terminal devices 101, 102, 103 are hardware, they can be various electronic devices with a display screen, including but not limited to smartphones, tablets, laptop computers, desktop computers, etc.; when terminal devices 101, 102, 103 are software, they can be installed in the above-listed electronic devices, and can be implemented as multiple software or software modules, or can be implemented as a single software or software module, which is not specifically limited here. When server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or can be implemented as a single server; when the server is software, it can be implemented as multiple software or software modules, or can be implemented as a single software or software module, which is not specifically limited here. For example, server 105 can be exemplified as a server that can provide a "cloud phone" service for terminal devices 101, 102, 103, that is, users can communicate with server 105 through terminal devices 101, 102, 103, and implement the "cloud phone" service with the "computing power" of the cloud provided by server 105. Correspondingly, for the convenience of understanding, in this article, a server that can provide the cloud computing power of a "cloud phone" (for example, server 105) is directly referred to as a "cloud phone".

[0026] Server 105 can provide various services through various built-in applications. Taking the cloud service-based file playback application as an example, server 105 can be used as a "cloud phone" to achieve the following effects through the cloud service-based file playback application: First, server 105 can respond when receiving file description information sent from a first device (for example, one of terminal devices 101, 102, 103) via network 104, and generate a corresponding file to be played based on the file description information; then, server 105 responds to the communication initiation request received from the first device and establishes a communication connection with a second device (for example, another one of terminal devices 101, 102, 103 that is different from the above first device); finally, server 105 plays the file to be played for the second device through a virtual input component, where the virtual input device is set in the cloud phone through virtualization technology.

[0027] Since the method for playing files in the present disclosure depends on and is applied to cloud phones, correspondingly, the device for playing files is generally also set in the server 105 exemplified as a "cloud phone".

[0028] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in

[0029] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 2 Figure 2 Please refer to

[0030] which is a flowchart of a process for playing files provided by an embodiment of the present disclosure, including process 200, which is applied to a cloud phone such as server 105 for example.

[0031] Step 201: In response to receiving file description information sent by a first device, generate a corresponding file to be played based on the file description information;

[0032] In an embodiment of the present disclosure, the execution entity for executing and implementing the method for playing files, that is, the "cloud phone" (for example, Figure 1 the server 105 shown), can pre-establish a communication connection with the first device used by the user (for example, one of the terminal devices 101, 102, 103). Then, the user can use such a first device to send file description information to the execution entity (that is, the cloud phone).

[0033] The file description information can generally be information expressed in forms such as voice, text, etc., and the content is expressed in forms such as natural language, specific code language, etc.

[0034] The file description information indicates the specific situation of the file to be played that needs to be provided by the cloud phone. For example, the file description information can indicate that the cloud phone needs to provide a video, audio, image, or text, etc., as well as the specific content of the video, audio, image, or text (for example, what kind of content the video includes), and the specific parameters of the file to be played (for example, the duration, resolution, etc. of the video).

[0035] Correspondingly, after receiving such file description information, the execution entity can generate a corresponding file to be played based on the file description information. For example, for a file to be played in the form of a video file, the execution entity can use a pre-configured generative model locally to generate a video file form of the file to be played that conforms to the specific parameters and has the content of the video based on the content of the video and the specific parameters indicated by the file description information.

[0036] ​Exemplarily, the file description information can be used to generate an introduction video for product XX with a playing time of YY seconds. The introduction video needs to introduce the characteristics of the product in three aspects: Z, C, and V, and the resolution of the introduction video is P.

[0037] Correspondingly, after receiving such file description information, the execution entity can correspondingly generate a video file T with a resolution of P and a content of introducing the characteristics of the product in three aspects: Z, C, and V for YY seconds as the file to be played (for example, the explanation video file for product XX).

[0038] In some optional implementation manners of this embodiment, the file description information sent by the first device can be sent in the form of a Shell instruction. That is, in this step, the execution entity can further and actually respond to the received file description information sent by the first device based on the Shell instruction to generate a corresponding file to be played based on the file description information.

[0039] A Shell instruction is a command used to interact with the command line interface (CLI) of an operating system. Shell is a user interface that allows users to communicate with the operating system by entering commands. By using Shell instructions, or in the case where using Shell instructions for interaction is allowed, users can use Shell to efficiently and accurately control the execution entity and provide file description information to the execution entity.

[0040] In some optional implementation manners of this embodiment, in order to enhance the ability of the cloud phone to generate files to be played, the execution entity can call a large language model (LLM) to generate a corresponding file to be played based on the file description information during the process of generating a corresponding file to be played based on the file description information. Thus, the powerful generation ability of the large language model can be utilized to generate a file to be played with higher quality (for example, a file to be played with richer content and better reading effect can be generated).

[0041] LLM is an artificial intelligence model designed to understand and generate human language, and LLM can perform corresponding processing operations based on the content it understands to obtain corresponding processing results. For example, after obtaining the file description information, LLM can be called according to a prompt such as "generate a corresponding file to be played based on the file description information" to generate a corresponding file to be played according to the content included in the file description information.

[0042] In practice, LLMs can be trained on large amounts of text data and can perform a wide range of tasks, including text summarization, translation, sentiment analysis, and so on. LLMs are characterized by their large scale, usually including a large number of parameters to help them learn complex patterns in language data. These models are typically based on deep learning architectures such as transformers, which help them provide better processing performance on various NLP tasks.

[0043] In addition, for generative models, at least some of the "prompt words" can be omitted by default configuration. For example, for the task of generating a corresponding file to be played based on file description information, the "prompt words" can be at least partially abbreviated based on the default configuration. For example, abbreviate "generate the corresponding file to be played". Accordingly, in such a case, the LLM can directly and quickly understand that it is necessary to generate a corresponding file to be played based on the incoming "file description information" as the "prompt word". Thus, the LLM can be called more efficiently.

[0044] In some alternative embodiments of the present disclosure, in order to call the LLM more effectively, the execution entity can also preprocess the file description information to further optimize the "prompt word". Thus, by improving the quality of the "prompt word", the "understanding cost" of the LLM can be reduced.

[0045] Specifically, during the process of calling the LLM, the execution entity can, in response to receiving the file description information sent by the first device, first perform word segmentation on the file description information to obtain a word segmentation result. For example, the execution entity can use methods such as the N-gram model and dictionary method to perform word segmentation on the file description information to obtain the corresponding word segmentation result.

[0046] Then, based on the entities hit by the word segmentation result, the execution entity generates a prompt word for indicating the actions of the large language model. For example, such entities can be "keywords" corresponding to the content, such as "video" and "audio" corresponding to the type of the file to be played, and "resolution P" corresponding to specific parameters. Accordingly, the execution entity can use these determined "entities" as "prompt words" to call the LLM to achieve the purpose of generating a corresponding file to be played based on the file description information.

[0047] Thus, the execution entity can filter out some information with low reference value through the processing of the file description information, so that the prompt word can more concisely and accurately reflect the user's "file generation requirements". Furthermore, through such word segmentation and filtering methods, the LLM can more clearly and accurately understand the content that the file to be played it needs to generate should have, enabling the LLM to generate a file to be played with higher quality.

[0048] Step 202: In response to receiving a communication initiation request from a first device, establish a communication connection with a second device corresponding to the communication initiation request.

[0049] In an embodiment of the present disclosure, after the execution entity generates a file to be played based on the above step 201, it may wait for a communication initiation request from the first device, or rather, a communication indication. Correspondingly, if the execution entity receives a communication initiation request sent by the first device, the execution entity may respond thereto, determine the second device indicated in the communication initiation request, and establish a communication connection with the second device corresponding to the communication initiation request.

[0050] Thereby, after the execution entity generates a corresponding file to be played for the user according to the user's needs (locally), it can be provided and shared with other users according to the user's instructions. This enables the user using the first device not to have to generate the file to be played locally on the first device and upload it before being able to use the cloud phone to share content and communicate with users using other devices (for example, the second device).

[0051] Correspondingly, such an interaction method also enables the cloud phone as the execution entity not to rely too much on the communication ability between the cloud phone and the first device when sharing and providing the file to be played according to the instructions of the user using the first device (for example, at this time, it only requires the bandwidth and communication ability to communicate the file description information between the first device and the cloud phone, rather than requiring the bandwidth and network resources to necessarily meet the transmission requirements of the file to be played).

[0052] In some optional implementation manners of this embodiment, in the process of establishing a communication connection with the second device corresponding to the communication initiation request in response to receiving the communication initiation request from the first device, the execution entity may also, in response to receiving the communication initiation request from the first device, first send a communication connection request to the second device corresponding to the communication initiation request to "inquire" the user of the second device.

[0053] Then, if the second device returns a confirmation message for the communication connection request, the execution entity may respond thereto to establish a communication connection with the second device. Thereby, it is avoided that the user of the second device is established with a communication connection with the cloud phone when not expecting to communicate or receive the file to be played, causing trouble to the user of the second device.

[0054] Step 203: Play the file to be played for the second device through the virtual input component.

[0055] In an embodiment of the present disclosure, after the execution entity establishes a communication connection with the second device based on the above step 202, it can provide and play these files to be played (for example, play the image frame sequence in the video file as the "image stream captured by the virtual camera") by indicating the recipient and viewer of the file to be played and by invoking the virtual input device pre-set in the execution entity (i.e., the cloud phone) through virtualization technology. The virtual input device may include a virtual camera and a virtual microphone.

[0056] In some alternative implementation manners of this embodiment, if the file to be played is a video file, when the execution entity plays the file to be played for the second device through the virtual input component, it may first write the image frame sequence of the video file into the virtual camera and write the audio frame sequence of the video file into the virtual microphone. Then, the video file is played for the second device through the virtual camera and the virtual microphone. For example, the virtual camera may be Snap Camera, ManyCam, XSplit VCam, etc. The virtual microphone may be Voicemeeter, Audio Repeater, etc.

[0057] Specifically, that is, when playing the video file, the execution entity may choose to separate the image frame sequence and the audio frame sequence from the video file. Then, the execution entity writes the image frame sequence as the "image stream" into the virtual camera and writes the audio frame sequence as the "audio stream" into the virtual microphone.

[0058] Then, the execution entity plays the video file for the second device through the virtual camera and the virtual microphone. That is, as discussed above, the execution entity may instruct the second device to invoke the virtual camera and the virtual microphone to obtain the streams "captured" therein. Or rather, the execution entity may instruct the second device to invoke the virtual camera and the virtual microphone and use the virtual camera and the virtual microphone as the "data source" to provide the "image stream" and the "audio stream" corresponding to the video file to the second device, so as to play the video file locally on the second device.

[0059] In this way, the second device can directly obtain the content of the video file by invoking the virtual input component, without the need to pre-configure and enter a specific resource library, nor to perform a large amount of configuration on the playback environment to meet the playback requirements. Thus, the efficiency of the cloud phone in providing content can be improved, and the configuration requirements of the cloud phone for the second device can be reduced.

[0060] In some alternative implementation manners of this embodiment, if the file to be played is an audio file, when the execution entity plays the file to be played for the second device through the virtual input component, the execution entity may first write the audio frame sequence of the audio file into the virtual microphone.

[0061] Then, play the audio file for the second device through the virtual microphone. Thus, for a file to be played in the form of an audio file that only includes audio content, the execution entity can also provide it to the second device by using the virtual microphone.

[0062] Thus, by differentially invoking different "virtual input components", different and personalized requirements for transferring files to be played by users can be met. For example, in a "video session" established between the first device and the second device using a real camera, the execution entity can generate an audio file according to the instruction of the first device, and use the communication connection between the execution entity and the second device to generate an audio file that meets the requirements of the file description information. Then, according to the instruction of the first device, during the "video session", use the cloud phone to play and provide the "audio file" to the second device as a substitute for the "audio file" to achieve differential intercommunication and content provision requirements.

[0063] For another example, in the "video session", the image stream and the video stream can be replaced simultaneously, so that the second device can directly obtain and play the file to be played in the window and interface of the "video session".

[0064] It should be understood that in some embodiments, based on different requirements, for the above-mentioned "writing process" (for example, writing an image frame sequence, an audio frame sequence), it can be executed after establishing a communication connection with the second device as discussed here (for example, in this case, making the virtual input component only after establishing a communication connection with the second device can avoid unexpected occupation of the resources of the virtual input component by premature writing), or it can be completed before establishing a communication connection with the second device corresponding to the communication initiation request (for example, in this case, making the virtual input component can be directly called by the second device after establishing a communication connection with the second device, which can improve the provision efficiency).

[0065] The method for playing a file provided by the embodiments of the present disclosure generates a corresponding file to be played based on the received file description information in response to receiving the file description information sent by the first device; establishes a communication connection with the second device corresponding to the communication initiation request in response to receiving the communication initiation request of the first device; plays the file to be played for the second device through the virtual input component, where the virtual input device is set in the cloud phone through virtualization technology. Thus, it can not only reduce the file provision cost when users use the cloud phone, but also reduce the communication requirements between the terminal device actually used by the users and the cloud phone, enabling users to provide and publish content more conveniently and at a lower cost by using the cloud phone.

[0066] In some embodiments, in order to improve the generation efficiency of the file to be played and save certain generation resources, in embodiments involving generating images such as videos, the execution entity may achieve this purpose by reusing and calling a pre-configured image frame sequence.

[0067] For example, in some embodiments, an image frame sequence library may be pre-maintained, such that after the execution entity generates a new image sequence frame, it stores the image sequence frame in the image frame sequence library, so that it can be "reused" by other generation processes later. Similarly, the image frame sequence library can also be used to store image frame sequences uploaded by users, for example, through a first device and a second device.

[0068] For the process of generating the file to be played in such a case, please further refer to Figure 3 . Figure 3 FIG. is a flowchart of a process for generating a file to be played provided by an embodiment of the present disclosure, which includes process 300. For example, process 300 may be an alternative or replacement implementation of step 201 in the above process 200.

[0069] Process 300 specifically includes the following steps:

[0070] Step 301: In response to receiving file description information sent by a first device, determine an image frame sequence generation task and an audio frame sequence generation task based on the file description information;

[0071] Specifically, the execution entity may, in response to receiving the file description information sent by the first device, first determine the image frame sequence generation task and the audio frame sequence generation task based on the file description information.

[0072] The image frame sequence generation task corresponds to the generation requirements related to the image frame sequence (taking a video file as an example, the image frame sequence generation task may correspond to the resolution of the video, the content that should be included in each image frame, etc.).

[0073] The audio frame sequence generation task corresponds to the generation requirements related to the audio frame sequence (for example, the "text content" of the audio in the video, etc.).

[0074] Then, for the image frame sequence generation task, the execution entity detects whether there is a historical image frame sequence in the pre-maintained image frame sequence library that can meet the image frame sequence generation task. For example, the execution entity may determine whether it can meet the image frame sequence generation task based on the semantics of the description information corresponding to each historical image frame sequence, the length of the historical image frame sequence, whether it can meet the requirements of the image frame sequence generation task, etc.

[0075] In some embodiments, corresponding comparison value ranges may be determined in advance for each comparison dimension (e.g., content dimension, time length dimension, resolution dimension, etc.). Then, the execution entity determines the comparison value corresponding to the dimension based on the "degree of conformity" under each comparison dimension.

[0076] Then, the execution entity determines whether the historical image frame sequence can meet the requirements of the image frame sequence generation task based on the direct addition or weighted addition result of the comparison values corresponding to each historical image frame sequence of a historical image frame sequence.

[0077] If there is a historical image frame sequence in the pre-maintained image frame sequence library that can meet the requirements of the image frame sequence generation task, the execution entity can, in response thereto, further execute step 302.

[0078] Step 302: Extract the historical image frame sequence from the image frame sequence library;

[0079] Step 303: Update the audio frame sequence generation task based on the historical image frame sequence;

[0080] Specifically, after determining the historical image frame sequence, the execution entity can update the audio frame sequence generation task based on the historical image frame sequence. For example, adjust the length and playback time of each audio frame in the audio frame sequence so that the audio sequence frames generated based on the updated audio frame sequence can "match" the historical image frame sequence. Thus, it is possible to avoid the situation where the matching degree between the audio frame sequence generated by the audio frame sequence generation task and the historical image frame sequence decreases due to the difference between the historical image frame sequence and the previously determined video frame sequence generation task.

[0081] Step 304: Generate an audio frame sequence based on the updated audio frame sequence generation task;

[0082] Specifically, in this step, the execution entity can, as discussed above, similar to the "description" of the audio frame sequence in the file description information, select to generate an audio frame sequence based on the updated audio frame sequence generation task, which will not be repeated here.

[0083] Step 305: Combine the historical image frame sequence and the audio frame sequence to generate a file description information to generate a corresponding file to be played.

[0084] Specifically, in this step, the execution entity can combine the historical image frame sequence determined in step 302 above and the audio frame sequence generated in step 304 to generate a file to be played in the form of a video file including an image track (i.e., the image frame sequence) and an audio track (i.e., the audio frame sequence).

[0085] In some embodiments, after generating the file to be played, the execution entity may also choose to provide it to the first device first to determine whether the generated file to be played can meet the needs of the user of the first device. For example, the execution entity may also play the file to be played for the first device through a virtual input component. For example, the execution entity may also instruct the first device to obtain the file to be played generated by the execution entity by calling a virtual microphone and a virtual camera.

[0086] Subsequently, if the user of the first device expects to modify the file to be played, they can return modification description information to the execution entity through the first device.

[0087] Accordingly, if the execution entity receives the modification description information returned by the first device for the file to be played, the execution entity can respond by modifying the file to be played based on the modification description information to obtain a modified file to be played. For example, the execution entity can also call the LLM, use the "modification description information" as a prompt, and use the LLM to modify the file to be played according to the modification description information.

[0088] Accordingly, in such a case, during the process of the execution entity playing the file to be played for the second device through the virtual input component, what is actually finally played for the second device through the virtual input component is the finally modified file to be played.

[0089] Thus, the execution entity can interact with the first device to allow the first device to further adjust the file to be played according to the needs, so that the file to be played can better meet the generation needs of the user of the first device.

[0090] It should be understood that in such an embodiment, "modification" can actually be in multiple rounds (for example, after the execution entity completes one round of modification, it continues to provide the modification result to the first device. Accordingly, the user of the first device can further return modification description information again for the next round of modification). And, in such an embodiment, if the user of the first device is satisfied with the content provided by the execution entity, they can also return a confirmation instruction to the execution entity accordingly, so that the execution entity can actually complete the generation process of the file to be played and wait for the "first device to send a communication initiation request" as discussed above.

[0091] In some embodiments, after the file to be played is generated, the user of the first device can also choose to configure permission requirements for the file to be played, so that other users who meet the permission requirements (e.g., the users of the third device) can "independently" obtain the file to be played, rather than having to wait for the instruction of the first device to obtain the file to be played from the cloud phone. Thus, such a method can not only reduce the instruction cost of the user of the first device (e.g., there is no need to initiate a request through communication every time), but also enable the file to be played to be more widely spread under the requirements of the user of the first device.

[0092] For example, in such a case, the third device can "actively" send a play request for the file to be played to the execution entity. Correspondingly, if the execution entity receives the play request for the file to be played from the third device, the execution entity can respond thereto and obtain the permission requirements configured by the first device for the file to be played.

[0093] In practice, the permission requirements can indicate the specific other devices that can be allowed by the user of the first device to obtain the file to be played, or can also indicate the characteristics of the other devices that can be allowed by the user of the first device to obtain the file to be played.

[0094] Correspondingly, if the execution entity determines that the play permission of the third device can meet the permission requirements, the execution entity can choose to establish a communication connection with the third device. Then, as discussed above, play the file to be played for the third device through the virtual input component.

[0095] For better understanding, the present disclosure also gives a specific implementation solution in combination with a specific application scenario. For this, please refer to Figure 4 . Figure 4 FIG. is a flowchart of the process of playing a file in an application scenario provided by an embodiment of the present disclosure, which includes process 400.

[0096] For ease of understanding, the discussion of process 400 will simultaneously refer to Figure 1 the architecture 100 shown for illustration. For example, in process 400, the server 105 can act as a "cloud phone" to provide services for the terminal devices 101 and 103 used by users (not shown in the figure).

[0097] First, the user can execute S401 through the terminal device 101 to send the file description information 410 to the server 105.

[0098] Accordingly, for the file description information 410, the server 105 may execute S402 to call the LLM to process the file description information 410, so as to generate a corresponding file to be played (for example, the video file 420) according to the user's requirements. And after the server 105 completes the generation of the video file 420, it may wait for a further instruction from the terminal device 101 (that is, wait for a communication initiation request sent by the terminal device 101).

[0099] Subsequently, if the terminal device 101 sends a communication initiation request to the server 105 by executing S403 (exemplarily, in this scenario, the communication initiation request corresponds to the terminal device 103), the server 105 may respond thereto and execute S404 to establish a communication connection with the terminal device 103.

[0100] Then, the server 105 continues to execute S405 to write the image frame sequence 421 of the video file 420 into the virtual camera 431 and write the audio frame sequence 422 of the video file 420 into the virtual microphone 432.

[0101] After completing S405, the server 105 may execute S406 to instruct the terminal device 103 to call the virtual camera 431 and the virtual microphone 432, so as to play the video file 420 for the terminal device 103 through the virtual input component.

[0102] Further referring to Figure 5 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a device for playing a file. This device embodiment corresponds to Figure 2 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0103] As Figure 5 shown, the device 500 for playing a file in this embodiment may include: a file to be played generation unit 501, a first communication connection establishment unit 502, and a file to be played playback unit 503. Among them, the file to be played generation unit 501 is configured to, in response to receiving file description information sent by a first device, generate a corresponding file to be played based on the file description information; the first communication connection establishment unit 502 is configured to, in response to receiving a communication initiation request from the first device, establish a communication connection with a second device corresponding to the communication initiation request; the file to be played playback unit 503 is configured to play the file to be played for the second device through a virtual input component, where the virtual input device is set in a cloud phone through virtualization technology.

[0104] In this embodiment, in the device 500 for playing files: For the specific processing of the to-be-played file generation unit 501, the first communication connection establishment unit 502, and the to-be-played file playing unit 503 and the technical effects brought thereby, reference can be respectively made to Figure 2 the relevant descriptions of steps 201-203 in the corresponding embodiment, which will not be elaborated herein.

[0105] In some optional implementation manners of this embodiment, the to-be-played file includes a video file. The to-be-played file playing unit 503 includes: a video file writing subunit configured to write the image frame sequence of the video file into a virtual camera and write the audio frame sequence of the video file into a virtual microphone; and a video file playing subunit configured to play the video file for a second device through the virtual camera and the virtual microphone.

[0106] In some optional implementation manners of this embodiment, the to-be-played file includes an audio file. The to-be-played file playing unit 503 includes: an audio file writing subunit configured to write the audio frame sequence of the audio file into a virtual microphone; and an audio file playing subunit configured to play the audio file for a second device through the virtual microphone.

[0107] In some optional implementation manners of this embodiment, the to-be-played file generation unit 501 includes: a generation task determination subunit configured to, in response to receiving file description information sent by a first device, determine an image frame sequence generation task and an audio frame sequence generation task based on the file description information; an image frame sequence extraction subunit configured to, in response to the existence of a historical image frame sequence in a pre-maintained image frame sequence library that can satisfy the image frame sequence generation task, extract the historical image frame sequence from the image frame sequence library; an audio frame generation task update subunit configured to update the audio frame sequence generation task based on the historical image frame sequence; an audio frame sequence generation subunit configured to generate an audio frame sequence based on the updated audio frame sequence generation task; and a to-be-played file generation subunit configured to combine the historical image frame sequence and the audio frame sequence to generate a to-be-played file corresponding to the file description information.

[0108] In some optional implementation manners of this embodiment, the to-be-played file generation unit 501 is further configured to, in response to receiving file description information sent by the first device based on a shell instruction, generate a corresponding to-be-played file based on the file description information.

[0109] In some alternative implementation manners of this embodiment, the first communication connection establishment unit 502 includes: a communication connection request subunit, configured to send a communication connection request to a second device corresponding to the communication initiation request in response to receiving the communication initiation request from the first device; and a first communication connection establishment subunit, configured to establish a communication connection with the second device in response to the second device returning confirmation information for the communication connection request.

[0110] In some alternative implementation manners of this embodiment, the apparatus 500 further includes: a to-be-played file return unit, configured to play a to-be-played file for the first device through a virtual input component; a to-be-played file modification unit, configured to modify the to-be-played file based on the received modification description information returned by the first device for the to-be-played file to obtain a modified to-be-played file; and the to-be-played file playing unit 503 is further configured to play the modified to-be-played file for the second device through the virtual input component.

[0111] In some alternative implementation manners of this embodiment, the apparatus 500 further includes: a permission requirement acquisition unit, configured to acquire the permission requirements configured by the first device for the to-be-played file in response to receiving a play request for the to-be-played file from a third device; a second communication connection establishment unit, configured to establish a communication connection with the third device in response to the play permission of the third device being able to meet the permission requirements; and the to-be-played file playing unit 503 is further configured to play the to-be-played file for the third device through the virtual input component.

[0112] In some alternative implementation manners of this embodiment, the to-be-played file generation unit 501 is further configured to generate a corresponding to-be-played file based on the received file description information sent by the first device by invoking a large language model.

[0113] In some alternative implementation manners of this embodiment, the to-be-played file generation unit 501 includes: a description information word segmentation subunit, configured to perform word segmentation processing on the file description information in response to receiving the file description information sent by the first device to obtain a word segmentation result; a prompt word generation subunit, configured to generate a prompt word for instructing the actions of the large language model based on the entities hit by the word segmentation result; and a large language model invocation subunit, configured to use the prompt word to invoke the large language model to generate a corresponding to-be-played file based on the file description information.

[0114] This embodiment exists as a device embodiment corresponding to the above method embodiment. The device for playing files provided in this embodiment generates a corresponding file to be played based on the file description information in response to receiving the file description information sent by the first device; establishes a communication connection with the second device corresponding to the communication initiation request in response to receiving the communication initiation request sent by the first device; and plays the file to be played for the second device through a virtual input component, where the virtual input device is set in the cloud phone through virtualization technology. Thus, it can not only reduce the file provision cost when users use the cloud phone, but also reduce the communication requirements between the terminal device actually used by the users and the cloud phone, enabling users to provide and publish content more conveniently and at a lower cost using the cloud phone.

[0115] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0116] Figure 6 FIG. shows a schematic block diagram of an exemplary electronic device 600 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0117] As Figure 6 shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0118] Multiple components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0119] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the method of playing a file. For example, in some embodiments, the method of playing a file can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method of playing a file described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the method of playing a file in any other suitable manner (e.g., by means of firmware).

[0120] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0121] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0122] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0123] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0124] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0125] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system to address the defects of high management difficulty and weak business scalability existing in traditional physical hosts and virtual private server (VPS) services. The server may also be a server of a distributed system or a server combined with a blockchain.

[0126] According to the technical solution of an embodiment of the present disclosure, in response to receiving file description information sent by a first device, a corresponding file to be played is generated based on the file description information; in response to receiving a communication initiation request from the first device, a communication connection with a second device corresponding to the communication initiation request is established; and the file to be played is played for the second device through a virtual input component, wherein the virtual input device is set in a cloud phone through virtualization technology. Thus, not only can the file provision cost be reduced when a user uses a cloud phone, but also the communication requirements between the terminal device actually used by the user and the cloud phone can be reduced, enabling the user to more conveniently and at a lower cost use the cloud phone to provide and publish content.

[0127] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution provided by the present disclosure can be achieved, and no limitation is made herein.

[0128] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A method for playing files, applied to a cloud phone, comprising: In response to receiving the file description information sent by the first device, generating a corresponding file to be played based on the file description information; In response to receiving the communication initiation request of the first device, establishing a communication connection with a second device corresponding to the communication initiation request; The file to be played is played for the second device through a virtual input component, wherein the virtual input device is set in the cloud phone through virtualization technology.

2. The method according to claim 1, wherein: The to-be-played file includes a video file, and playing the to-be-played file for the second device through the virtual input component includes: Writing the image frame sequence of the video file into a virtual camera, and writing the audio frame sequence of the video file into a virtual microphone; The video file is played for the second device through the virtual camera and the virtual microphone.

3. The method according to claim 1, wherein: The to-be-played file includes an audio file, and playing the to-be-played file for the second device through the virtual input component includes: Writing the audio frame sequence of the audio file into a virtual microphone; The audio file is played for the second device through the virtual microphone.

4. The method according to claim 1, wherein: The to-be-played file includes a video file, and in response to receiving the file description information sent by the first device, generating a corresponding to-be-played file based on the file description information includes: In response to receiving the file description information sent by the first device, determining the image frame sequence generation task and the audio frame sequence generation task based on the file description information; In response to the presence of a historical image frame sequence that can satisfy the image frame sequence generation task in a pre-maintained image frame sequence library, extracting the historical image frame sequence from the image frame sequence library; Based on the historical image frame sequence, updating the audio frame sequence generation task; Based on the updated audio frame sequence generation task, generate an audio frame sequence; The historical image frame sequence and the audio frame sequence are combined to generate the file description information to generate a corresponding file to be played.

5. The method according to claim 1, wherein: The step of generating a corresponding to-be-played file based on the file description information in response to receiving the file description information sent by the first device includes: In response to receiving the file description information sent by the first device based on the shell instruction, a corresponding file to be played is generated based on the file description information.

6. The method according to claim 1, wherein: The step of establishing a communication connection with a second device corresponding to the communication initiation request in response to receiving the communication initiation request from the first device includes: In response to receiving a communication initiation request from the first device, sending a communication connection request to a second device corresponding to the communication initiation request; In response to the second device returning confirmation information in response to the communication connection request, a communication connection with the second device is established.

7. The method according to claim 1, further comprising: Playing the to-be-played file for the first device through the virtual input component; In response to receiving the modification description information returned by the first device for the file to be played, modifying the file to be played based on the modification description information to obtain a modified file to be played; as well as The step of playing the to-be-played file for the second device through the virtual input component includes: The modified file to be played is played for the second device through the virtual input component.

8. The method according to claim 1, further comprising: In response to receiving a play request from a third device for the file to be played, obtaining a permission requirement configured by the first device for the file to be played; In response to the playback permission of the third device satisfying the permission requirement, establishing a communication connection with the third device; The file to be played is played for the third device through the virtual input component.

9. The method according to any one of claims 1 to 8, wherein: The step of generating a corresponding to-be-played file based on the file description information in response to receiving the file description information sent by the first device includes: In response to receiving the file description information sent by the first device, a corresponding to-be-played file is generated based on the file description information by calling a large language model.

10. The method according to claim 9, wherein: In response to receiving the file description information sent by the first device, generating a corresponding to-be-played file based on the file description information by calling the large language model, including: In response to receiving the file description information sent by the first device, processing the file description information by word segmentation to obtain a word segmentation result; Based on the entity hit by the word segmentation result, generating a prompt word for indicating the action of the large language model; The prompt word is used to call a large language model to generate a corresponding to-be-played file based on the file description information.

11. A device for playing files, applied to a cloud phone, comprising: a to-be-played file generating unit, configured to generate a corresponding to-be-played file based on the file description information in response to receiving the file description information sent by the first device; A first communication connection establishing unit, configured to, in response to receiving a communication initiation request from the first device, establish a communication connection with a second device corresponding to the communication initiation request; The to-be-played file playing unit is configured to play the to-be-played file for the second device through a virtual input component, wherein the virtual input device is set in the cloud phone through virtualization technology.

12. The device according to claim 11, wherein The to-be-played file includes a video file, and the to-be-played file playing unit includes: A video file writing subunit is configured to write the image frame sequence of the video file into the virtual camera, and write the audio frame sequence of the video file into the virtual microphone; The video file playing subunit is configured to play the video file for the second device through the virtual camera and the virtual microphone.

13. The device according to claim 11, wherein: The to-be-played file includes an audio file, and the to-be-played file playing unit includes: an audio file writing subunit, configured to write the audio frame sequence of the audio file into the virtual microphone; The audio file playing subunit is configured to play the audio file for the second device through the virtual microphone.

14. The device according to claim 11, wherein: The to-be-played file generating unit comprises: a generation task determination subunit, configured to, in response to receiving the file description information sent by the first device, determine the image frame sequence generation task and the audio frame sequence generation task based on the file description information; an image frame sequence extraction subunit, configured to extract the historical image frame sequence from the image frame sequence library in response to the presence of the historical image frame sequence that can meet the image frame sequence generation task in the pre-maintained image frame sequence library; an audio frame generation task updating subunit, configured to update the audio frame sequence generation task based on the historical image frame sequence; an audio frame sequence generation subunit, configured to generate an audio frame sequence based on the updated audio frame sequence generation task; The to-be-played file generating subunit is configured to combine the historical image frame sequence and the audio frame sequence, generate the file description information and generate the corresponding to-be-played file.

15. The device according to claim 11, wherein The to-be-played file generating unit is further configured to, in response to receiving file description information sent by the first device based on the shell instruction, generate a corresponding to-be-played file based on the file description information.

16. The device according to claim 11, wherein The first communication connection establishing unit includes: a communication connection request subunit, configured to, in response to receiving a communication initiation request from a first device, send a communication connection request to a second device corresponding to the communication initiation request; The first communication connection establishing subunit establishes a communication connection with the second device in response to the second device returning confirmation information in response to the communication connection request.

17. The apparatus according to claim 11, further comprising: a to-be-played file returning unit, configured to play the to-be-played file for the first device through the virtual input component; a to-be-played file modification unit, configured to, in response to receiving modification description information returned by the first device for the to-be-played file, modify the to-be-played file based on the modification description information to obtain a modified to-be-played file; as well as The to-be-played file playing unit is further configured to play the modified to-be-played file for the second device through a virtual input component.

18. The apparatus according to claim 11, further comprising: a permission requirement obtaining unit, configured to obtain the permission requirement configured by the first device for the file to be played in response to receiving a play request for the file to be played by a third device; The second communication connection establishing unit is configured to establish a communication connection with the third device in response to the playback permission of the third device satisfying the permission requirement; and the to-be-played file playing unit is also configured to play the to-be-played file for the third device through the virtual input component.

19. The device according to any one of claims 11 to 18, wherein: The to-be-played file generating unit is further configured to, in response to receiving the file description information sent by the first device, generate a corresponding to-be-played file based on the file description information by calling a large language model.

20. The device according to claim 19, wherein The to-be-played file generating unit comprises: The description information segmentation subunit is configured to, in response to receiving the file description information sent by the first device, segment the file description information to obtain a segmentation result; A prompt word generation subunit is configured to generate a prompt word for indicating an action of the large language model based on the entity hit by the word segmentation result; The large language model calling subunit is configured to use the prompt word to call the large language model to generate a corresponding to-be-played file based on the file description information.

21. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for playing a file according to any one of claims 1 to 10.

22. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method for playing a file according to any one of claims 1 to 10.

23. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method for playing a file according to any one of claims 1 to 10.