Method, apparatus, device and medium for obtaining live video
By identifying the user's sign language recognition function in live videos and converting voice into sign language information, the problem of hearing-impaired users obtaining information is solved and the user experience is improved.
Patent Information
- Application Number
- CN202211348863.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-31
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-10-31
AI Technical Summary
It is difficult for hearing-impaired users to obtain effective information from live videos.
By identifying whether the user is turning on the sign language recognition function and converting the voice information into sign language information in the synthesized video in the connected state, a video with sign language information is generated and pushed to the user.
It improves the information acquisition ability of hearing-impaired users when watching live videos and enhances the user experience.
Smart Images

Figure CN115695849B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical fields of natural language processing, cloud computing, big data, computer vision, speech technology, deep learning, etc. in artificial intelligence, and particularly relates to a method, apparatus, device, and medium for obtaining live video. Background Art
[0002] Currently, with the continuous development of video technology, video live streaming has gradually become the main way of information transmission. Users can obtain product information, news information, course information, etc. by watching video live streams.
[0003] However, for hearing-impaired users, it is difficult for hearing-impaired users to obtain effective information from live videos. Therefore, there is an urgent need for a method for obtaining live videos so that hearing-impaired users can obtain the information in live videos in a timely and accurate manner. Summary of the Invention
[0004] The present disclosure provides a method, apparatus, device, and medium for obtaining live video that enables hearing-impaired users to accurately obtain effective information in live videos.
[0005] According to a first aspect of the present disclosure, there is provided a method for obtaining live video, including:
[0006] In response to a viewing request initiated by a first user, obtaining a first video indicated by the viewing request; wherein, the viewing request includes a sign language identifier, and the sign language identifier indicates whether the sign language recognition function is enabled; the first video has a status identifier, and the status identifier indicates whether the first video is a synthesized video in a co-hosting state;
[0007] If it is determined that the sign language identifier indicates that the sign language recognition function is enabled and the status identifier indicates that the first video is a synthesized video in a co-hosting state, then converting the voice information in the first video into sign language information to obtain a first video with sign language information, and pushing the first video with sign language information to the first user.
[0008] According to a second aspect of the present disclosure, there is provided an apparatus for obtaining live video, including:
[0009] A first obtaining unit, configured to obtain a first video indicated by a viewing request in response to a viewing request initiated by a first user; wherein, the viewing request includes a sign language identifier, and the sign language identifier indicates whether the sign language recognition function is enabled; the first video has a status identifier, and the status identifier indicates whether the first video is a synthesized video in a co-hosting state;
[0010] A first conversion unit, configured to convert the voice information in the first video into sign language information to obtain a first video with sign language information if it is determined that the sign language identifier represents enabling the sign language recognition function and the status identifier represents that the first video is a combined video in a co-hosting state;
[0011] A first pushing unit, configured to push the first video with sign language information to the first user.
[0012] According to a third aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method as described in the first aspect.
[0013] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method as described in the first aspect.
[0014] According to a fifth aspect of the present disclosure, there is provided a computer program product, including: a computer program, the computer program is stored in a readable storage medium, and at least one processor of an electronic device can read the computer program from the readable storage medium, and when the at least one processor executes the computer program, the electronic device is caused to execute the method as described in the first aspect.
[0015] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. Description of the Drawings
[0016] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0017] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure;
[0018] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure;
[0019] Figure 3 is a schematic diagram of a scenario provided by the present disclosure;
[0020] Figure 4 is a schematic diagram according to the third embodiment of the present disclosure;
[0021] Figure 5 is a schematic diagram according to the fourth embodiment of the present disclosure;
[0022] Figure 6 is a schematic diagram according to the fifth embodiment of the present disclosure;
[0023] Figure 7 is a schematic diagram according to the sixth embodiment of the present disclosure;
[0024] Figure 8 is a block diagram of an electronic device for implementing the method for obtaining a live video according to an embodiment of the present disclosure. Detailed implementation manners
[0025] The following makes an explanation of exemplary embodiments of the present disclosure with reference to the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0026] Currently, with the continuous development of video technology, more and more users can obtain commodity information, news information, course information, etc. by watching live videos.
[0027] When conducting a video live broadcast, the voice of the anchor when speaking is the main way to convey information. However, for hearing-impaired users, it is difficult to obtain effective information. Therefore, a way to obtain a live video is needed so that hearing-impaired users can obtain the information in the live video.
[0028] To avoid at least one of the above technical problems, the inventors of the present disclosure have obtained the inventive concept of the present disclosure through creative labor: when a first user initiates a viewing request, first obtain the first video indicated by the viewing request. If it is determined that the sign language identifier carried in the viewing request indicates that the first user has enabled the sign language recognition function and, based on the status identifier of the first video, it is determined that the first video is a synthesized video in a co-hosting state, then convert the voice information in the first video into sign language information to obtain a first video with sign language information, and push the first video with sign language information to the first user so that the first user can obtain the information in the first video.
[0029] According to the above inventive concept, the present disclosure provides a method, apparatus, device, and medium for obtaining a live video, which are applied to technical fields such as natural language processing, cloud computing, big data, computer vision, speech technology, and deep learning in artificial intelligence, so that hearing-impaired users can obtain the information in the live video when watching the live video, improving the user experience.
[0030] In the technical solution of the present disclosure, the processing of the user's personal information, such as collection, storage, use, processing, transmission, provision, and disclosure, complies with the provisions of relevant laws and regulations and does not violate public order and good customs.
[0031] Figure 1 It is a schematic diagram according to the first embodiment of the present disclosure. As Figure 1 shown, the method for obtaining a live video according to an embodiment of the present disclosure includes:
[0032] S101. In response to a viewing request initiated by a first user, obtain a first video indicated by the viewing request; wherein, the viewing request includes a sign language identifier, and the sign language identifier represents whether the sign language recognition function is enabled; the first video has a status identifier, and the status identifier represents whether the first video is a composite video in a co-hosting state.
[0033] Exemplarily, the execution subject of this embodiment may be a device for obtaining a live video. The obtaining device may be a server (such as a cloud server or a local server), a computer, a terminal device, a processor, a chip, etc., and this embodiment is not limited.
[0034] In this embodiment, the viewing request initiated by the first user includes: a sign language identifier. Wherein, the sign language identifier may represent whether the first user currently enables the sign language recognition function.
[0035] After receiving the viewing request initiated by the user, first, the first video indicated by the viewing request will be obtained, that is, the video of the live broadcast room that the first user hopes to watch. And, the obtained first video carries a status identifier, where the status identifier can be used to represent whether the currently obtained first video is a composite video in a co-hosting state. It can be understood that when the host accessed by the first user is in a co-hosting, the first video obtained by the first user is a composite video of the video of the host accessed by the first user and the videos of the other hosts co-hosting with the host.
[0036] In one example, when the user enters the live video application software, information on whether to enable the sign language recognition function can be pushed to the user; or, the sign language recognition function can be automatically enabled for the user based on the user's historical usage habits, so that the user can obtain a live video with sign language information.
[0037] S102. If it is determined that the sign language identifier represents that the sign language recognition function is enabled and the status identifier represents that the first video is a composite video in a co-hosting state, convert the voice information in the first video into sign language information to obtain a first video with sign language information, and push the first video with sign language information to the first user.
[0038] Exemplarily, when it is determined that the sign language identifier in the viewing request of the first user indicates that the first user has enabled the sign language recognition function, then the first user is considered a hearing-impaired user at this time. At this time, if it is further determined that the status identifier carried by the first video indicates that the current first video is a composite video in a co-hosting state, then the voice information in the first video can be converted to obtain the sign language information corresponding to the voice information, and then the first video carrying the sign language information is obtained.
[0039] In one example, when converting the voice information in the first video into sign language information, all the voice information included in the first video can be converted into text information at this time, and then the corresponding sign language information can be obtained based on the obtained text information.
[0040] In one example, when pushing the first video with sign language information to the user, at this time, the first video and the sign language information are displayed on the same layer, that is, multiple display areas are divided in the same layer, where one display area is used to display the sign language information, and the remaining display areas are used to display the video images corresponding to each host in the co-hosting host of the first video respectively.
[0041] It can be understood that in this embodiment, when the first user is a hearing-impaired user, the first user can enable the sign language recognition function at this time. When the first video requested by the first user is a composite video in a co-hosting state, the voice information in the first video can be converted into sign language information at this time, so that the first user can accurately obtain the effective information in the first video in the co-hosting state and improve the viewing experience of the first user.
[0042] To enable readers to more deeply understand the implementation principle of the present disclosure, the following Figure 2 is Figure 1 shown in the embodiment is further refined.
[0043] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure. As Figure 2 shown, the method for obtaining a live video according to the embodiment of the present disclosure includes:
[0044] S201. In response to a viewing request initiated by a first user, obtain a first video indicated by the viewing request; wherein, the viewing request includes a sign language identifier, and the sign language identifier indicates whether the sign language recognition function is enabled; the first video has a status identifier, and the status identifier indicates whether the first video is a composite video in a co-hosting state.
[0045] Exemplarily, the execution subject of this embodiment may be a live video acquisition device, and the acquisition device may be a server (such as a cloud server or a local server), a computer, a terminal device, a processor, a chip, etc., which is not limited in this embodiment.
[0046] For the specific principle of step S201, reference can be made to step S101, which will not be elaborated here.
[0047] S202. If it is determined that the sign language identifier represents enabling the sign language recognition function, and the status identifier represents the composite video of the first video in the co-hosting state, and it is determined that there is a non-hearing-impaired host among the co-hosts corresponding to the first video, then the first video is split to obtain at least one second video and at least one third video, where the second video is the video uploaded by the hearing-impaired host; the third video is the video uploaded by the non-hearing-impaired host.
[0048] Exemplarily, in this embodiment, when it is determined that the first user has enabled the sign language recognition function and the obtained first video is a composite video in the co-hosting state, at this time, it can be further determined whether there is a non-hearing-impaired host among the co-hosts included in the first video. When it is determined that there is a non-hearing-impaired host in the first video, when determining the sign language information corresponding to the voice information in the first video, at this time, the first video information is split to obtain at least one second video and at least one third video. Among them, the second video can be understood as the live video of the live broadcast room corresponding to the hearing-impaired user among the co-hosts included in the first video. The third video can be understood as the live video of the live broadcast room of the non-hearing-impaired host among the co-hosts included in the first video.
[0049] In one example, when determining whether there is a non-hearing-impaired host in the first video, if the sign language of the hearing-impaired host is translated into corresponding audio information during the live broadcast and the audio information is generated by voice with a preset tone, at this time, it can be detected whether there is a sound in the voice information in the first video that is the same as the preset tone. If so, it indicates that there is a hearing-impaired host; if not, it indicates that there is a non-hearing-impaired host. Or, in a possible implementation manner, since the first video is synthesized from the videos of multiple co-hosts, it can be detected whether each co-host has enabled the sign language recognition function. If enabled, it indicates that there is a hearing-impaired host; if not enabled, it indicates that there is a non-hearing-impaired host.
[0050] In one example, the step of "determining that there is a non-hearing-impaired host among the co-hosts corresponding to the first video" in step S202 may include the following steps: "obtaining the lip movement information of the co-hosts in the first video; performing voice conversion processing on the lip movement information of the co-hosts to obtain the converted voice information; if it is determined that the converted voice information is consistent with the voice information of the first video, then determining that there is a non-hearing-impaired host among the co-hosts corresponding to the first video".
[0051] Exemplarily, in this embodiment, when determining whether there is a non-hearing-impaired host in the first video, the lip movements of the co-host in the first video can be analyzed. That is, the lip movement information corresponding to the co-host can be processed through voice conversion to obtain the converted voice information corresponding to the lip movement information. Then, the converted voice information is compared with the voice information in the first video. If the two are consistent, it can be considered that the co-host is a non-hearing-impaired host. If the converted voice information is inconsistent with the voice information in the first video, it can be considered that the co-host is a hearing-impaired host. Among them, when comparing the converted voice information with the voice information in the first video, if the similarity between the two is greater than a preset threshold, it can be considered that the two are consistent.
[0052] It can be understood that in this embodiment, the method of determining whether there is a non-hearing-impaired host in the first video by comparing the voice information converted based on the lip movement information of the co-host in the first video and the original voice information in the first video is beneficial to improving the accuracy of the judgment result, and further ensures that the first user can accurately obtain the effective information in the first video.
[0053] S203. Perform sign language translation processing on the voice information in the third video to obtain a gesture video.
[0054] Exemplarily, in this embodiment, after determining the third video in the first video, since the first user has enabled the sign language recognition function and the third video in the first video is a video uploaded by a non-hearing-impaired user, the third video can be subjected to sign language translation processing to obtain a gesture video. Among them, the gesture video can be understood as the gesture video obtained after the voice information in the third video is translated into sign language in real time. For example, for a third video, a gesture video can be generated based on all the voice information in the third video, that is, one third video corresponds to one gesture video.
[0055] In one example, step S203 includes the following steps:
[0056] The first step of step S203: Identify the voice information corresponding to each non-hearing-impaired host in the third video.
[0057] Exemplarily, in this embodiment, when determining the gesture video corresponding to the third video, since there may be multiple live users in the video corresponding to a live room, after obtaining the third video, the voice information corresponding to each non-hearing-impaired host in the third video will also be determined. For example, when performing this step, the lip movement recognition method can also be used to determine the voice information corresponding to each non-hearing-impaired host.
[0058] In one example, the first step of step S203 may include the following steps: "Filter the speech information in the third video according to the preset audio information in the preset database to obtain the filtered speech information; perform sign language translation processing on the filtered speech information to obtain a gesture video."
[0059] Exemplarily, the preset database contains at least one preset audio information. For example, the preset audio information includes live special effect audio information, pop music audio information, the host's historical live audio information, etc.
[0060] In practical applications, since during the host's live broadcast, in addition to the speaking voice of the live broadcast, the live video usually also includes other sounds. To facilitate the first user to accurately obtain the host's speaking voice when watching the video, in this embodiment, when identifying the speech information corresponding to the non-hearing-impaired host in the third video, first, the speech information in the third video can be filtered based on the preset audio information in the preset database. For example, according to the live feature audio or pop music in the preset database, the live special effect audio and pop music in the third video can be filtered out. Another example is that the preset database can include the audio information in the host's historical live videos, and then based on the audio information in the historical live videos, determine the voice characteristics of the host's voice, and then the speech information that does not correspond to the host's voice can be filtered out in the speech information of the third video to obtain the filtered speech information. After that, in the filtered speech information, identify the speech information corresponding to each non-hearing-impaired host.
[0061] It can be understood that in this embodiment, by filtering the speech information in the third video, it is beneficial to accurately obtain the speech information corresponding to the non-hearing-impaired host, and further make the gesture video obtained by sign language translation processing more accurate, which is beneficial for the first user to obtain the effective information in the live video.
[0062] The second step of step S203: Perform sign language translation processing on the speech information corresponding one by one to the non-hearing-impaired hosts in the third video to obtain the gesture videos corresponding one by one to the non-hearing-impaired hosts in the third video.
[0063] Exemplarily, after obtaining the speech information corresponding to each non-hearing-impaired host in the third video, sign language translation processing can be performed on the speech information corresponding to each non-hearing-impaired host respectively, and then the gesture videos corresponding one by one to each non-hearing-impaired host can be obtained. After that, after merging the obtained gesture videos of the hearing-impaired hosts, the second video, and the third video, the first video with sign language information can be obtained.
[0064] It can be understood that when determining the gesture video corresponding to the third video in this embodiment, a gesture video can be generated for each non-hearing-impaired anchor in the third video, so that when the first user watches the first video, the first user can determine the information corresponding to each different anchor through different gesture videos, which is beneficial to improving the viewing experience of the first user.
[0065] In a possible implementation manner, in order to facilitate the first user to determine the non-hearing-impaired anchor corresponding to the gesture translation video, the same identifier can be set for the display screen of the non-hearing-impaired anchor and the display screen of the gesture video of the non-hearing-impaired anchor. For example, the same label can be used, or a rectangular frame of the same color can be used for annotation, etc. The specific implementation manner is not specifically limited in this application.
[0066] In a possible implementation manner, the first user can also select the generation method of the gesture video he hopes. For example, one gesture video can correspond to multiple non-hearing-impaired anchors, or one gesture video can correspond to one non-hearing-impaired anchor.
[0067] In one example, on the basis of the above example, after the second step of the above step S203, the following steps are further included: obtaining the user image features corresponding to each non-hearing-impaired anchor in the third video; based on the user image features corresponding to each non-hearing-impaired anchor in the third video, adjusting the image features of the virtual person in the gesture video corresponding to the third video to obtain the adjusted gesture video.
[0068] Exemplarily, in this embodiment, after generating the gesture video corresponding to each non-hearing-impaired anchor in the third video, the user image features corresponding to each non-hearing-impaired anchor will also be obtained. For example, the user image features may include the user's clothing features, the user's hairstyle features, the user's accessory (such as bracelets, glasses, earrings, etc.) features. Furthermore, after obtaining the above user image features, the gesture video of the non-hearing-impaired anchor in the third video corresponding to the user image features can be adjusted. That is, based on the user image features, the image features of the digital person in the gesture video are adjusted to be the same as the user image features. After that, after merging and processing the adjusted gesture videos, the second video, and the third video, the first video with sign language information can be obtained. The digital person in the gesture video can be understood as a digital virtual person generated by combining artificial intelligence technology.
[0069] In practical applications, when converting voice information into a gesture video, the voice information can be converted into text information at this time, and then the obtained text information is converted into corresponding sign language instructions, and then based on the obtained sign language instructions, the digital person is controlled to change the gestures of the digital person according to the corresponding sign language instructions.
[0070] It can be understood that in this embodiment, the digital human in the gesture video corresponding to the non-hearing-impaired anchor can be adjusted in terms of image features based on the user characteristics of the non-hearing-impaired anchor, so that the first user can determine the non-hearing-impaired anchor corresponding to each gesture video, which is beneficial to improving the user viewing experience.
[0071] S204. Merge the second video, the third video, and the gesture video to obtain a first video with sign language information, and push the first video with sign language information to the first user.
[0072] Exemplarily, in this embodiment, after obtaining the gesture video corresponding to the third video, the second video (i.e., the video uploaded by the hearing-impaired user in the first video), the third video, and the gesture video can be merged, and then a first video with sign language information can be obtained. After that, the first video with sign language information is displayed to the user in real time. In practical applications, when performing the merging process, the merging can be performed according to the time axis corresponding to each video at this time.
[0073] It can be understood that when it is determined that there is a non-hearing-impaired anchor among the co-hosts of the first video, at this time, the first video will be split to obtain the videos uploaded by the live rooms corresponding to each co-host. The split videos include at least one second video and at least one third video. And when performing sign language translation processing, only the third video needs to be translated, and there is no need to translate the voice information in the second video uploaded by the hearing-impaired user. Since in the related art, when a hearing-impaired user uploads a video, the sign language of the hearing-impaired user is converted into corresponding voice information and then uploaded, therefore, the method provided in this embodiment can reduce the processing time during sign language translation processing and improve the efficiency of generating the first video with sign language information.
[0074] S205. In response to the moving operation of the first user, control the preset playback window to move to the position indicated by the moving operation; the display area corresponding to the preset playback window is used to display the gesture video in the first video with sign language information.
[0075] Exemplarily, in this embodiment, after the user obtains the first video with sign language information, the gesture video included in the video can be displayed in the display area corresponding to the preset playback window. And under the moving operation of the user, the user can adjust the position of the preset playback window in the user display screen. For example, the gesture video corresponding to each non-hearing-impaired anchor can be displayed in the lower right corner of the display area of its corresponding third video. The first user can move the window corresponding to the gesture video, thereby changing its position on the current display screen.
[0076] It can be understood that the preset playback window corresponding to the gesture video of the non-hearing-impaired user in this embodiment can be moved according to the user's movement operation, so that the display screen in the current display can conform to the viewing habits of the first user, which is beneficial to improving the user's viewing experience.
[0077] In addition, in a possible implementation manner, the size of the preset playback window can also be adjusted under the user's operation.
[0078] S206. In response to the connection request initiated by the first user, obtain the fourth video of the first user; the connection request includes: a sign language identifier.
[0079] Exemplarily, when the first user needs to connect, the connection request initiated by the first user may also carry a sign language identifier. And when the first user initiates a connection, the fourth video of the first user will be obtained in real time, that is, the image and voice of the first user will be obtained in real time.
[0080] S207. If it is determined that the sign language identifier indicates that the sign language recognition function has been enabled, convert the sign language information of the fourth video into voice information to obtain a fourth video with voice information.
[0081] Exemplarily, when it is determined that the sign language identifier in the connection request of the first user indicates that the user has enabled the sign language recognition function, at this time, it is determined that the first user is a hearing-impaired user. Then, the sign language information in the obtained fourth video can be translated into corresponding voice information, and then a fourth video with voice information can be generated, so that the other connection hosts who connect with the first user can accurately obtain the effective information in the live video of the first user, and the other users who watch the connection video of the first user can also accurately understand the effective information output by the first user.
[0082] In this embodiment, when it is determined that there is a non-deaf-mute host among the co-hosts of the first video, at this time, the first video will be split to obtain the videos uploaded by the live rooms corresponding to each co-host. The split videos include at least one second video and at least one third video. Moreover, when performing sign language translation processing, only the third video needs to be translated, and there is no need to translate the voice information in the second video uploaded by the deaf-mute user. Since in the related art, when a deaf-mute user uploads a video, the sign language of the deaf-mute user will be converted into voice information and then uploaded after obtaining the corresponding voice information, therefore, the method provided in this embodiment can reduce the processing time during sign language translation processing and improve the efficiency of generating the first video with sign language information. In addition, when determining the sign language video corresponding to the third video, a sign language video can be generated for each non-deaf-mute host in the third video, so that when the first user watches the first video, different information corresponding to different hosts can be determined through different sign language videos, which is beneficial to improving the viewing experience of the first user.
[0083] Figure 3 is a schematic diagram of a scenario provided by the present disclosure. As Figure 3 shown in the figure, in the figure, the cloud server is used to cache the videos uploaded by the hosts. The push stream engine and the pull stream engine can be regarded as devices in each terminal device (for example, mobile terminals such as computers and mobile phones used by users). In addition, the audio recognition engine, the translation engine, the sign language recognition engine, and the voice conversion engine can be regarded as devices in the devices of the deaf-mute user side, or can also be devices other than the devices of the deaf-mute user side, and no specific limitation is made in this embodiment.
[0084] Among them, the push stream engine and the pull stream engine can communicate with the server through RTP (Real-time Transport Protocol) or RTN (Real-time Network).
[0085] Other users can also encapsulate the videos to be transmitted into a streaming media format (such as the flv format) for transmission through the http (Hyper Text Transfer) protocol, or can also adopt the RTMP (Real Time Messaging Protocol) protocol, or can also adopt the HLS (HTTP Live Streaming) adaptive bitrate streaming media transmission protocol based on the http protocol for transmission.
[0086] On the hearing-impaired user side, the hearing-impaired user can enable the sign language recognition function. When the hearing-impaired user's device 301 can send a viewing request to the first video pulling engine 302, at this time, the first video pulling engine 302 on the hearing-impaired user side can obtain the original video indicated by the viewing request from the server 303. After that, when there are non-hearing-impaired users in the obtained original video, at this time, the first video pulling engine 302 can split out the videos corresponding to the non-hearing-impaired users from the original video, and send the split videos to the speech recognition engine 304. The speech recognition engine 304 can convert the speech information in the split videos into text information, and then send the text information to the translation engine 305. The translation engine 305 can convert the text information into corresponding sign language instructions, and then control the digital human actions based on the obtained sign language instructions to generate gesture videos. After that, after the first video pulling engine 302 merges the original video and the gesture videos, it sends them to the hearing-impaired user's device 301. Among them, the original video and the gesture videos can be displayed in different windows on the hearing-impaired user side. For example, the original video can be displayed in the display area of the preset main window, and the gesture video can be displayed in the display area of the preset secondary window.
[0087] In addition, when the hearing-impaired user initiates a co-hosting while the sign language recognition function is enabled, at this time, the co-hosting video collected by the hearing-impaired user's device 301 can be synchronized to the first video pushing engine 306, and then sent by the first video pushing engine 306 to the sign language recognition engine 307, so that the sign language recognition engine 307 can convert the sign language of the hearing-impaired user into corresponding text, and send the converted text to the speech conversion engine 308. The speech conversion engine 308 can convert the received text into corresponding speech information and send it back to the first video pushing engine 306. After that, the first video pushing engine 306 will merge the original co-hosting video and the converted speech information and then upload them to the server.
[0088] The non-hearing-impaired user device 309 that co-hosts with the hearing-impaired user can obtain the video with speech information uploaded by the hearing-impaired host through its corresponding second video pulling engine 310, so that the non-hearing-impaired user can know the information that the hearing-impaired user needs to express. And when the non-hearing-impaired user uploads a live video, it can also be directly uploaded to the server by its corresponding second video pushing engine 311 without translation processing.
[0089] Figure 4 It is a schematic diagram according to the third embodiment of the present disclosure. As Figure 4 shown, the method for obtaining a live video according to the embodiment of the present disclosure includes:
[0090] S401. In response to a viewing request initiated by a first user, obtain a first video indicated by the viewing request. The viewing request includes a sign language identifier, which indicates whether the sign language recognition function is enabled. The first video has a status identifier, which indicates whether the first video is a composite video in a co-hosting state.
[0091] Exemplarily, the execution subject of this embodiment may be a device for obtaining live videos. The device may be a server (such as a cloud server or a local server), a computer, a terminal device, a processor, a chip, etc. This embodiment does not make any limitations.
[0092] S402. If it is determined that the sign language identifier indicates that the sign language recognition function is enabled and the status identifier indicates that the first video is a composite video in a co-hosting state, convert the voice information in the first video into sign language information to obtain a first video with sign language information, and push the first video with sign language information to the first user.
[0093] Exemplarily, the specific principles of steps S401 and S402 can be referred to steps S101 and S102, which will not be elaborated here.
[0094] S403. If it is determined that the sign language identifier indicates that the sign language recognition function is enabled, the status identifier indicates that the first video is a composite video in a co-hosting state, and there is no non-hearing-impaired host among the co-hosting anchors corresponding to the first video, push the first video to the first user.
[0095] Exemplarily, in this embodiment, when the sign language identifier in the viewing request of the first user indicates that the user has enabled the sign language recognition function, the first video is a composite video in a co-hosting state, and there is no non-hearing-impaired host among the co-hosting anchors of the first video, then in the scenario where all the anchors in the co-hosting video are hearing-impaired hosts, the first video can be directly pushed to the first user at this time.
[0096] It can be understood that when the first user enables the sign language recognition function, although the first user will be considered a hearing-impaired user, since all the anchors in the first video are hearing-impaired hosts, the first video can be directly pushed to the first user, avoiding the phenomenon of wasting processing resources by only determining that the first video needs to be processed for sign language translation based on the fact that the first user has enabled the sign language recognition function identifier.
[0097] S404. If it is determined that the sign language identifier indicates that the sign language recognition function is enabled, the status identifier indicates that the first video is not a composite video in a co-hosting state, and the second user is a non-hearing-impaired host, convert the voice information in the first video into sign language information to obtain a first video with sign language information, and push the first video with sign language information to the first user. The second user is the anchor in the first video.
[0098] Exemplarily, in this embodiment, when the first user enables the sign language recognition function, but the first video is not a synthesized video in the co-hosting state, it is still necessary to determine whether the host in the first video (i.e., the second user) is a hearing-impaired host. Specifically, when determining whether the second user is a hearing-impaired host, the method in the above embodiment can also be used and will not be elaborated in this embodiment.
[0099] When it is determined that the second user is not a hearing-impaired host, at this time, the voice information in the first video can be translated into corresponding sign language information to obtain the first video with sign language information.
[0100] It should be noted that in the above scenario, if there are also multiple non-hearing-impaired hosts in the first video, at this time, the voice information corresponding to each non-hearing-impaired host can also be recognized first, and the sign language information corresponding to the voice information of each non-hearing-impaired host can be generated. The specific implementation method can adopt the method in the above embodiment and will not be elaborated in this embodiment.
[0101] It can be understood that in this embodiment, when the first user enables the sign language recognition function, if the first video is a non-co-hosting video, at this time, it can also be determined whether the second user in the first video is a hearing-impaired user. If the user is not a hearing-impaired user, sign language translation is required so that the first user can obtain the effective information in the first video.
[0102] S405. If it is determined that the sign language identifier indicates the enabling of the sign language recognition function, and the status identifier indicates that the first video is not a synthesized video in the co-hosting state, and the second user is a hearing-impaired host, then the first video is pushed to the first user.
[0103] Exemplarily, in this embodiment, when the first user enables the sign language recognition function, but the first video is not a synthesized video in the co-hosting state, and in addition, it is determined that the second user in the first video is a hearing-impaired host, at this time, there is no need to perform sign language translation processing on the voice information in the first video, and the first video can be directly pushed to the first user.
[0104] It can be understood that in this implementation, when the first user enables the sign language recognition function, at this time, if it is determined that the first video is not a synthesized video in the co-hosting state, and the host in the first video is a hearing-impaired host, at this time, there is no need to perform additional processing on the first video and it can be directly pushed, thereby avoiding the time-consuming of additional translation processing and reducing the occupancy rate of processing resources.
[0105] S406. If it is determined that the sign language identifier indicates that the sign language recognition function is not enabled, then the first video is pushed to the first user.
[0106] Exemplarily, in this embodiment, when it is determined that the sign language identifier in the viewing request of the first user indicates that the sign language recognition function is not enabled, the first user is considered a non-hearing-impaired user. At this time, the first video is directly pushed to the first user so that the first user can directly watch it. It can be understood that in this embodiment, by recognizing the sign language identifier in the viewing request, if the user does not enable the sign language recognition function, the first video can be directly pushed without additional processing, reducing the occupation of processing resources.
[0107] Figure 5 is a schematic diagram according to the fourth embodiment of the present disclosure, as Figure 5 shown, the live video acquisition device 500 of the embodiment of the present disclosure includes:
[0108] A first acquisition unit 501, configured to acquire a first video indicated by a viewing request in response to a viewing request initiated by a first user; wherein, the viewing request includes a sign language identifier, and the sign language identifier represents whether the sign language recognition function is enabled; the first video has a status identifier, and the status identifier represents whether the first video is a synthesized video in a co-hosting state.
[0109] A first conversion unit 502, configured to convert the voice information in the first video into sign language information to obtain a first video with sign language information if it is determined that the sign language identifier indicates that the sign language recognition function is enabled and the status identifier indicates that the first video is a synthesized video in a co-hosting state;
[0110] A first push unit 503, configured to push the first video with sign language information to the first user.
[0111] Exemplarily, the device in this embodiment can execute the technical solutions in the above method, and the specific implementation process and technical principle are the same, which will not be elaborated here.
[0112] Figure 6 is a schematic diagram according to the fifth embodiment of the present disclosure, as Figure 6 shown, the live video acquisition device 600 of the embodiment of the present disclosure includes:
[0113] A first acquisition unit 601, configured to acquire a first video indicated by a viewing request in response to a viewing request initiated by a first user; wherein, the viewing request includes a sign language identifier, and the sign language identifier represents whether the sign language recognition function is enabled; the first video has a status identifier, and the status identifier represents whether the first video is a synthesized video in a co-hosting state.
[0114] A first conversion unit 602, configured to convert the voice information in the first video into sign language information to obtain a first video with sign language information if it is determined that the sign language identifier indicates that the sign language recognition function is enabled and the status identifier indicates that the first video is a synthesized video in a co-hosting state;
[0115] The first pushing unit 603 is configured to push a first video with sign language information to a first user.
[0116] In one example, the first conversion unit 602 includes:
[0117] The splitting module 6021 is configured to, if it is determined that the sign language identifier indicates the activation of the sign language recognition function, and the status identifier indicates that the first video is a composite video in a co-hosting state, and it is determined that there is a non-hearing-impaired host among the co-hosting anchors corresponding to the first video, perform splitting processing on the first video to obtain at least one second video and at least one third video, where the second video is a video uploaded by a hearing-impaired host; the third video is a video uploaded by a non-hearing-impaired host.
[0118] The translation module 6022 is configured to perform sign language translation processing on the voice information in the third video to obtain a gesture video.
[0119] The merging module 6023 is configured to perform merging processing on the second video, the third video, and the gesture video to obtain a first video with sign language information.
[0120] In one example, the translation module 6022 includes:
[0121] The recognition sub-module is configured to recognize the voice information corresponding to each non-hearing-impaired host in the third video.
[0122] The translation sub-module is configured to perform sign language translation processing on the voice information corresponding to each non-hearing-impaired host in the third video to obtain the gesture video corresponding to each non-hearing-impaired host in the third video.
[0123] In one example, the apparatus further includes:
[0124] The first acquisition sub-module is configured to acquire the user image features corresponding to each non-hearing-impaired host in the third video.
[0125] The adjustment sub-module is configured to, based on the user image features corresponding to each non-hearing-impaired host in the third video, perform image feature adjustment on the digital human in the gesture video corresponding to each non-hearing-impaired host in the third video to obtain the adjusted gesture video corresponding to each non-hearing-impaired host in the third video.
[0126] In one example, the recognition sub-module is specifically configured to:
[0127] Perform filtering processing on the voice information in the third video according to the preset audio information in the preset database to obtain the filtered voice information; recognize the voice information corresponding to each non-hearing-impaired host in the third video from the filtered voice information.
[0128] In one example, the splitting module 6021 includes:
[0129] A second acquisition sub-module, configured to acquire the lip movement information of the co-host in the first video if it is determined that the sign language identifier represents the activation of the sign language recognition function and the status identifier represents that the first video is a synthesized video in a co-hosting state.
[0130] A conversion sub-module, configured to perform voice conversion processing on the lip movement information of the co-host to obtain the converted voice information.
[0131] A determination sub-module, configured to determine that there is a non-hearing-impaired host among the co-hosts corresponding to the first video if it is determined that the converted voice information is consistent with the voice information of the first video;
[0132] A splitting sub-module, configured to perform splitting processing on the first video to obtain at least one second video and at least one third video, where the second video is a video uploaded by a hearing-impaired host; the third video is a video uploaded by a non-hearing-impaired host.
[0133] In one example, the apparatus further includes:
[0134] A second pushing unit 604, configured to push the first video to the first user if it is determined that the sign language identifier represents the activation of the sign language recognition function, the status identifier represents that the first video is a synthesized video in a co-hosting state, and there is no non-hearing-impaired host among the co-hosts corresponding to the first video.
[0135] In one example, the apparatus further includes:
[0136] A second conversion unit 605, configured to convert the voice information in the first video into sign language information to obtain a first video with sign language information if it is determined that the sign language identifier represents the activation of the sign language recognition function, the status identifier represents that the first video is not a synthesized video in a co-hosting state, and the second user is a non-hearing-impaired host; the second user is the host in the first video.
[0137] A third pushing unit 606, configured to push the first video with sign language information to the first user.
[0138] In one example, the apparatus further includes:
[0139] A fourth pushing unit 607, configured to push the first video to the first user if it is determined that the sign language identifier represents the activation of the sign language recognition function, the status identifier represents that the first video is not a synthesized video in a co-hosting state, and the second user is a hearing-impaired host; the second user is the host in the first video.
[0140] In one example, the apparatus further includes:
[0141] The fifth pushing unit 608 is configured to push the first video to the first user if it is determined that the sign language identifier indicates that the sign language recognition function is not enabled.
[0142] In one example, the apparatus further includes:
[0143] The second obtaining unit 609 is configured to obtain a fourth video of the first user in response to a co-hosting request initiated by the first user; the co-hosting request includes: a sign language identifier.
[0144] The third conversion unit 610 is configured to convert the sign language information of the fourth video into voice information to obtain a fourth video with voice information if it is determined that the sign language identifier indicates that the sign language recognition function is enabled.
[0145] In one example, the apparatus further includes:
[0146] The control unit 611 is configured to control the preset playback window to move to the position indicated by the movement operation in response to the movement operation of the first user; the display area corresponding to the preset playback window is used to display the gesture video in the first video with sign language information.
[0147] Exemplarily, the apparatus in this embodiment may execute the technical solutions in the above method, and the specific implementation process and technical principle are the same, which will not be elaborated here.
[0148] Figure 7 It is a schematic diagram according to the sixth embodiment of the present disclosure. As Figure 7 shown, the electronic device 700 in the present disclosure may include: a processor 701 and a memory 702.
[0149] A memory 702 for storing programs; the memory 702 may include volatile memory, such as random-access memory (RAM), such as static random-access memory (SRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM), etc.; the memory may also include non-volatile memory, such as flash memory. The memory 702 is used to store computer programs (such as application programs and functional modules for implementing the above methods), computer instructions, etc. The above computer programs, computer instructions, etc. can be partitioned and stored in one or more memories 702. And the above computer programs, computer instructions, data, etc. can be called by the processor 701.
[0150] The above computer programs, computer instructions, etc. can be partitioned and stored in one or more memories 702. And the above computer programs, computer instructions, data, etc. can be called by the processor 701.
[0151] A processor 701 for executing the computer programs stored in the memory 702 to implement each step in the method involved in the above embodiments.
[0152] For details, please refer to the relevant descriptions in the foregoing method embodiments.
[0153] The processor 701 and the memory 702 may be of an independent structure or an integrated structure integrated together. When the processor 701 and the memory 702 are of an independent structure, the memory 702 and the processor 701 can be coupled and connected through a bus 703.
[0154] The electronic device in this embodiment can execute the technical solutions in the above method, and the specific implementation process and technical principle are the same, which will not be elaborated here.
[0155] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, which includes: a computer program, the computer program is stored in a readable storage medium, and at least one processor of the electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to enable the electronic device to execute the solution provided in any of the above embodiments.
[0156] Figure 8FIG. 0 shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0157] As Figure 8 shown, the device 800 includes a computing unit 801 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0158] A plurality of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0159] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the method for acquiring live video. For example, in some embodiments, the method for acquiring live video can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the method for acquiring live video described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the method for acquiring live video in any other suitable manner (e.g., by means of firmware).
[0160] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0161] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0162] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0163] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0164] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0165] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with a blockchain.
[0166] It should be understood that various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0167] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for obtaining a live video, comprising: In response to a viewing request initiated by a first user, obtaining a first video indicated by the viewing request; wherein, a sign language identifier is included in the viewing request, and the sign language identifier represents whether the sign language recognition function is enabled; the first video has a status identifier, and the status identifier represents whether the first video is a composite video in a co-hosting state; In the case of determining that the sign language identifier represents that the sign language recognition function is enabled and the status identifier represents that the first video is a composite video in a co-hosting state, obtaining the lip information of the co-hosting anchor in the first video; the composite video represents a composite video of the video of the anchor accessed by the first user and the videos of the other anchors co-hosting with the anchor; Performing voice conversion processing on the lip information of the co-hosting anchor to obtain converted voice information; If it is determined that the converted voice information is consistent with the voice information of the first video, then it is determined that there is a non-hearing-impaired anchor among the co-hosting anchors corresponding to the first video, and then the first video is split to obtain at least one second video and at least one third video, wherein the second video is a video uploaded by a hearing-impaired anchor; the third video is a video uploaded by a non-hearing-impaired anchor; Performing sign language translation processing on the voice information in the third video to obtain a gesture video; Performing merging processing on the second video, the third video, and the gesture video to obtain a first video with sign language information, and pushing the first video with sign language information to the first user; In response to a co-hosting request initiated by a first user, obtaining a fourth video of the first user; the co-hosting request includes: a sign language identifier; If it is determined that the sign language identifier represents that the sign language recognition function has been enabled, then converting the sign language information of the fourth video into voice information to obtain a fourth video with voice information.
2. The method according to claim 1, wherein Performing sign language translation processing on the voice information in the third video to obtain a gesture video, including: Identifying the voice information corresponding one by one to the non-hearing-impaired anchors in the third video; Performing sign language translation processing on the voice information corresponding one by one to the non-hearing-impaired anchors in the third video to obtain the gesture videos corresponding one by one to the non-hearing-impaired anchors in the third video.
3. The method according to claim 2, further comprising: Obtaining the user image features corresponding one by one to the non-hearing-impaired anchors in the third video; Based on the user image features corresponding one by one to the non-hearing-impaired anchors in the third video, adjusting the image features of the digital humans in the gesture videos corresponding one by one to the non-hearing-impaired anchors in the third video to obtain the adjusted gesture videos corresponding one by one to the non-hearing-impaired anchors in the third video.
4. The method according to any one of claims 2 or 3, wherein, Identifying the voice information corresponding one by one to the non-hearing-impaired anchors in the third video, including: Performing filtering processing on the voice information in the third video according to the preset audio information in the preset database to obtain filtered voice information; Identifying the voice information corresponding one by one to the non-hearing-impaired anchors in the third video from the filtered voice information.
5. The method according to any one of claims 1-3 further includes: If it is determined that the sign language identifier indicates that the sign language recognition function is enabled, and the status identifier indicates that the first video is a composite video in a co-hosting state, and there is no non-hearing-impaired host among the co-hosts corresponding to the first video, then push the first video to the first user.
6. The method according to any one of claims 1-3 further includes: If it is determined that the sign language identifier indicates that the sign language recognition function is enabled, and the status identifier indicates that the first video is not a composite video in a co-hosting state, and the second user is a non-hearing-impaired host, then convert the voice information in the first video into sign language information to obtain a first video with sign language information, and push the first video with sign language information to the first user; The second user is the host in the first video.
7. The method according to any one of claims 1-3 further includes: If it is determined that the sign language identifier indicates that the sign language recognition function is enabled, and the status identifier indicates that the first video is not a composite video in a co-hosting state, and the second user is a hearing-impaired host, then push the first video to the first user, where the second user is the host in the first video.
8. The method according to any one of claims 1-3 further includes: If it is determined that the sign language identifier indicates that the sign language recognition function is not enabled, then push the first video to the first user.
9. The method according to any one of claims 2-3 further includes: In response to the mobile operation of the first user, control the preset playback window to move to the position indicated by the mobile operation; The display area corresponding to the preset playback window is used to display the gesture video in the first video with sign language information.
10. An apparatus for obtaining a live video, including: A first obtaining unit, configured to obtain a first video indicated by the viewing request in response to a viewing request initiated by a first user; wherein, the viewing request includes a sign language identifier, and the sign language identifier indicates whether the sign language recognition function is enabled; the first video has a status identifier, and the status identifier indicates whether the first video is a composite video in a co-hosting state; A first conversion unit, configured to convert the voice information in the first video into sign language information to obtain a first video with sign language information if it is determined that the sign language identifier indicates that the sign language recognition function is enabled and the status identifier indicates that the first video is a composite video in a co-hosting state; A first pushing unit, configured to push the first video with sign language information to the first user; A second obtaining unit, configured to obtain a fourth video of the first user in response to a co-hosting request initiated by the first user; the co-hosting request includes: a sign language identifier; A third conversion unit, configured to convert the sign language information of the fourth video into voice information to obtain a fourth video with voice information if it is determined that the sign language identifier indicates that the sign language recognition function is enabled; The first conversion unit includes: A splitting module, configured to, if it is determined that the sign language identifier indicates the activation of the sign language recognition function, and the status identifier indicates that the first video is a composite video in a co-hosting state, and it is determined that there is a non-hearing-impaired host among the co-hosting hosts corresponding to the first video, perform splitting processing on the first video to obtain at least one second video and at least one third video, where the second video is a video uploaded by a hearing-impaired host; the third video is a video uploaded by a non-hearing-impaired host; A translation module, configured to perform sign language translation processing on the voice information in the third video to obtain a gesture video; A merging module, configured to perform merging processing on the second video, the third video, and the gesture video to obtain the first video with sign language information; The splitting module includes: A second acquisition sub-module, configured to, if it is determined that the sign language identifier indicates the activation of the sign language recognition function, and the status identifier indicates that the first video is a composite video in a co-hosting state, acquire the lip movement information of the co-hosting hosts in the first video; A conversion sub-module, configured to perform voice conversion processing on the lip movement information of the co-hosting hosts to obtain the converted voice information; A determination sub-module, configured to, if it is determined that the converted voice information is consistent with the voice information of the first video, determine that there is a non-hearing-impaired host among the co-hosting hosts corresponding to the first video; A splitting sub-module, configured to perform splitting processing on the first video to obtain at least one second video and at least one third video.
11. The apparatus according to claim 10, wherein, The translation module includes: An identification sub-module, configured to identify the voice information corresponding to the non-hearing-impaired host in the third video one by one; A translation sub-module, configured to perform sign language translation processing on the voice information corresponding to the non-hearing-impaired host in the third video one by one to obtain the gesture video corresponding to the non-hearing-impaired host in the third video one by one.
12. The apparatus according to claim 11, further comprising: A first acquisition sub-module, configured to acquire the user image features corresponding to the non-hearing-impaired host in the third video one by one; An adjustment sub-module, configured to, based on the user image features corresponding to the non-hearing-impaired host in the third video one by one, perform image feature adjustment on the digital human in the gesture video corresponding to the non-hearing-impaired host in the third video one by one to obtain the adjusted gesture video corresponding to the non-hearing-impaired host in the third video one by one.
13. The device according to any one of claims 11 or 12, wherein, The identification sub-module is specifically configured to: Perform filtering processing on the voice information in the third video according to the preset audio information in the preset database to obtain the filtered voice information; Identify the voice information corresponding to the non-hearing-impaired host in the third video one by one from the filtered voice information.
14. The apparatus according to any one of claims 10-12, further comprising: A second push unit, configured to, if it is determined that the language identifier indicates that the sign language recognition function has been activated, and the status identifier indicates that the first video is a composite video in a co-hosting state, and there is no non-hearing-impaired host among the co-hosting hosts corresponding to the first video, push the first video to the first user.
15. The apparatus according to any one of claims 10-12, further comprising: A second conversion unit, configured to convert the voice information in the first video into sign language information to obtain a first video with sign language information if it is determined that the sign language identifier indicates the activation of the sign language recognition function, the status identifier indicates that the first video is not a synthesized video in a co-hosting state, and the second user is a non-hearing-impaired host; The second user is the host in the first video; A third pushing unit, configured to push the first video with sign language information to the first user.
16. The apparatus according to any one of claims 10-12, further comprising: A fourth pushing unit, configured to push the first video to the first user if it is determined that the sign language identifier indicates the activation of the sign language recognition function, the status identifier indicates that the first video is not a synthesized video in a co-hosting state, and the second user is a hearing-impaired host, and the second user is the host in the first video.
17. The apparatus according to any one of claims 10-12, further comprising: A fifth pushing unit, configured to push the first video to the first user if it is determined that the sign language identifier does not indicate the activation of the sign language recognition function.
18. The apparatus according to any one of claims 10-12, further comprising: A control unit, configured to control a preset playback window to move to the position indicated by the movement operation in response to a movement operation of the first user; The display area corresponding to the preset playback window is used to display the gesture video in the first video with sign language information.
19. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-9.
20. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-9.
21. A computer program product, comprising a computer program, where the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1-9.
Citation Information
Patent Citations
Video uploading method and device, video playing method and device, client equipment and storage medium
CN110730360A
Microphone-connected live broadcast interaction method and device and computer equipment
CN113573083A