Cross-device information broadcasting method, apparatus, system, device, and storage medium
The cross-device information playback method processes user-end information on the server side, generates playback files, and displays multimedia resources on the playback terminal. This solves the problem of users securely obtaining information in different scenarios and achieves convenient and secure information access.
Patent Information
- Application Number
- CN202411954861.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2044-12-27
AI Technical Summary
In modern life, users find it difficult to safely access and consume information without affecting their main tasks in different scenarios, especially when their hands or attention are limited, such as while driving or exercising, where traditional visual reading methods pose safety risks.
Through the cross-device information playback method, the user sends information to the server for processing. The server segments the text and converts it into audio segments, and generates a playback file by combining multimedia resources. The playback terminal displays multimedia resources during playback, realizing the acquisition of information by combining voice and vision.
It enables users to quickly obtain information through voice and vision without affecting their primary tasks, improving the convenience and security of information acquisition and reducing security risks.
Smart Images

Figure CN119766791B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of Internet, and particularly relates to a cross-device information broadcasting method, device, system, equipment and storage medium. BACKGROUND
[0002] With the rapid development of mobile Internet, using smart terminals such as mobile phones, tablets and the like to obtain information instantly has become a part of daily life. However, in the fast-paced lifestyle of modern society, users often switch frequently between different life scenes, and in the scenes of driving, sports, housework and the like where hands or attention are limited, and other situations such as visual fatigue cannot be visually read, users hope to obtain and consume various information such as news information, articles and the like without affecting the main task under the premise that the traditional visual reading method is no longer applicable, and may also bring safety hazards in driving and other scenes that require high concentration of attention.
[0003] Therefore, there is an urgent need for an information obtaining method that can meet the dual needs of users for information obtaining convenience and safety. SUMMARY
[0004] Therefore, in order to solve the above technical problems, the present application provides a cross-device information broadcasting method, device, system, equipment and readable storage medium.
[0005] Specifically, the present application is realized by the following technical solutions:
[0006] According to a first aspect of an embodiment of the present application, a cross-device information broadcasting method is provided, applied to a server, and the method comprises:
[0007] receiving an indication message from a user terminal, and determining to-be-broadcast information according to the indication message; the indication message comprises multi-modal information for positioning the to-be-broadcast information;
[0008] segmenting text in the to-be-broadcast information according to the position of a multimedia resource in the to-be-broadcast information, to obtain a plurality of text segments;
[0009] converting the text segments into corresponding audio segments, and sequentially organizing the audio segments and the multimedia resource according to the position sequence of the text segments and the multimedia resource in the to-be-broadcast information to obtain a broadcasting file, so that a playing terminal accesses the server to obtain the broadcasting file, and the corresponding multimedia resource is displayed in the process of playing the broadcasting file.
[0010] Optionally, the segmenting the text in the to-be-read information to obtain a plurality of text segments comprises: traversing the to-be-read information, continuously collecting text until a multimedia resource is detected or the traversal ends in a case that text is detected, and taking the currently collected text as a text segment; and in a case that a multimedia resource is detected, associating the multimedia resource with the last collected text segment.
[0011] The sequentially organizing the audio segments and the multimedia resources to obtain the read file comprises: adding the associated multimedia resource at the playing end position of each audio segment, and integrating to obtain the read file.
[0012] Optionally, the method further comprises: for each multimedia resource, determining at least one text segment adjacent to the multimedia resource; obtaining the graphic-text relevance of each sentence in the multimedia resource and the adjacent text segment, and determining the relevant sentence with the highest graphic-text relevance therefrom.
[0013] The sequentially organizing the audio segments and the multimedia resources to obtain the read file comprises: determining the audio playing position of the relevant sentence for the audio segment corresponding to the adjacent text segment; adding the multimedia resource at the audio playing position of the relevant sentence corresponding to the multimedia resource, and integrating to obtain the read file.
[0014] Optionally, the sequentially organizing the audio segments and the multimedia resources to obtain the read file comprises:
[0015] A unique identifier is set for each text segment; the audio segment corresponding to the text segment, the added multimedia resource and the resource type thereof are organized into a data slice together with the unique identifier; and all data slices are sequentially integrated according to the position sequence of the text segments to form a complete read file.
[0016] Optionally, the sequentially organizing the audio segments and the multimedia resources to obtain the read file comprises:
[0017] In a case that the multimedia resource is a picture, description text of the picture is generated through a graphic-text network, an audio segment corresponding to the description text is obtained, and the audio segment corresponding to the text segment and the multimedia resource is sequentially organized to obtain the read file; or in a case that the multimedia resource is a video, an audio file in the video is extracted as an audio segment, and the audio segment corresponding to the text segment and the multimedia resource is sequentially organized to obtain the read file.
[0018] Optionally, the determining the to-be-read information according to the indication message comprises:
[0019] determining a content type of the indication message; the content type comprises any one of a webpage link, information text; parsing the indication message according to the content type to determine to-be-read information.
[0020] Optionally, the determining the content type of the indication message comprises:
[0021] acquiring a field value representing a message type in a transmission encapsulation format of the indication message; in a case where the field value represents a link, determining that the content type is a webpage link; in a case where the field value represents text, detecting, by using a regular expression, whether a string in the indication message comprises a link in a text form, and if yes, determining that the content type is a webpage link, and if no, determining that the content type is information text; in a case where the field value represents an image, extracting text information from the image, and detecting, by using a regular expression, whether the text information belongs to a webpage link or information text.
[0022] Optionally, the parsing the indication message according to the content type to determine to-be-read information comprises:
[0023] in a case where the content type is a webpage link, locating to webpage content according to the link in the indication message; parsing a webpage structure of the webpage content, and sequentially acquiring webpage text and multimedia resources as to-be-read information according to a relative position relationship between the webpage text and the multimedia resources;
[0024] or, in a case where the content type is information text, acquiring a field value representing message content in a transmission encapsulation format of the indication message as the to-be-read information.
[0025] Optionally, the server comprises a custom server on a WeChat platform service; the receiving the indication message from the user terminal comprises: receiving the indication message forwarded by a WeChat server; the WeChat server receives the indication message sent by a user through a client of the WeChat platform service.
[0026] According to a second aspect of an embodiment of the present application, a cross-device information reading device is provided, the device being applied to a server, and the device comprises:
[0027] an information acquisition module, configured to receive an indication message from a user terminal, and determine to-be-read information according to the indication message; the indication message comprises multi-modal information used for locating to-be-read information;
[0028] a text segmentation processing module, configured to perform segmentation processing on text in the to-be-read information according to a position of a multimedia resource in the to-be-read information, to obtain a plurality of text segments;
[0029] The broadcast file construction module is configured to convert the text segments into corresponding audio segments, and sequentially organize the audio segments and the multimedia resources according to the position sequence of the text segments and the multimedia resources in the to-be-broadcast information to obtain a broadcast file, so that the playing terminal accesses the server to obtain the broadcast file, and the corresponding multimedia resources are displayed in the process of playing the broadcast file.
[0030] Optionally, the text segment processing module is specifically configured to:
[0031] The text segment processing module is configured to: traverse the to-be-broadcast information, continuously collect text until a multimedia resource is detected or the end of the traversal, and take the currently collected text as a text segment; and in the case that a multimedia resource is detected, associate the multimedia resource with the last collected text segment.
[0032] The broadcast file construction module is specifically configured to: add the associated multimedia resource at the end of the playing of each audio segment, and integrate the broadcast file.
[0033] Optionally, the apparatus further comprises:
[0034] The server determines at least one text segment adjacent to each multimedia resource, and obtains the graphic-text relevance of each sentence in the multimedia resource and the adjacent text segments, and determines the relevant sentence with the highest graphic-text relevance.
[0035] The broadcast file construction module is specifically configured to: determine the audio playing position of the relevant sentence for the audio segment corresponding to the adjacent text segment; add the multimedia resource at the audio playing position of the relevant sentence corresponding to the multimedia resource, and integrate the broadcast file.
[0036] Optionally, the broadcast file construction module is specifically configured to:
[0037] A unique identifier is set for each text segment; the audio segment corresponding to the text segment, the added multimedia resource and the resource type of the multimedia resource are organized into a data slice with the unique identifier; and all data slices are sequentially integrated according to the position sequence of the text segments to form a complete broadcast file.
[0038] Optionally, the broadcast file construction module is specifically configured to:
[0039] In a case where the multimedia resource is a picture, description text of the picture is generated through a picture-text network, an audio segment corresponding to the description text is obtained, and a reading file is obtained by sequentially organizing the text segment and the audio segment corresponding to the multimedia resource; or in a case where the multimedia resource is a video, an audio file in the video is extracted as an audio segment, and a reading file is obtained by sequentially organizing the text segment and the audio segment corresponding to the multimedia resource.
[0040] Optionally, the information obtaining module is specifically configured to: determine a content type of the indication message; the content type includes any one of a webpage link and information text; and parse the indication message according to the content type to determine the information to be read.
[0041] Optionally, the information obtaining module includes the following when used to determine the content type of the indication message:
[0042] obtain a field value representing a message type in a transmission encapsulation format of the indication message;
[0043] in a case where the field value represents a link, determine that the content type is a webpage link;
[0044] in a case where the field value represents text, detect whether a string in the indication message includes a link in a text form through a regular expression, and if yes, determine that the content type is a webpage link, and if no, determine that the content type is information text;
[0045] in a case where the field value represents an image, extract text information from the image, and detect whether the text information belongs to a webpage link or information text through a regular expression.
[0046] Optionally, the information obtaining module includes the following when used to parse the indication message according to the content type to determine the information to be read:
[0047] in a case where the content type is a webpage link, locate to webpage content according to a link in the indication message; parse a webpage structure of the webpage content, and sequentially obtain webpage body text and a multimedia resource as the information to be read according to a relative position relationship between the webpage body text and the multimedia resource;
[0048] or, in a case where the content type is information text, obtain a field value representing message content in a transmission encapsulation format of the indication message as the information to be read.
[0049] Optionally, the server includes a custom server on a WeChat platform service; and the information obtaining module includes the following when used to receive the indication message from the user end:
[0050] The indication message is sent by a user through a client of the WeChat platform.
[0051] According to a third aspect of the embodiments of the present application, a cross-device information reading system is provided, which comprises a user terminal, a server and a playing terminal. The user terminal sends an indication message to the server, so that the server determines to-be-read information according to the indication message. The indication message comprises multi-modal information for locating the to-be-read information. The server segments text in the to-be-read information according to the position of multimedia resources in the to-be-read information to obtain a plurality of text segments, converts the text segments into corresponding audio segments, and sequentially organizes the audio segments and the multimedia resources according to the position sequence of the text segments and the multimedia resources in the to-be-read information to obtain a reading file. The playing terminal accesses the server to obtain the reading file and displays corresponding multimedia resources in the process of playing the reading file.
[0052] According to a fourth aspect of the embodiments of the present application, an electronic device is provided, which comprises a memory and a processor. The memory is used to store a computer program. The processor is used to execute the cross-device information reading method by calling the computer program.
[0053] According to a fifth aspect of the embodiments of the present application, a computer readable storage medium is provided, which stores a computer program. The program is executed by a processor to implement the cross-device information reading method.
[0054] In the above technical solution provided by the present application, the user sends information to be read to the server through the user terminal in the form of an indication message. The server determines to-be-read information by recognizing and analyzing the type of the indication message. Then, the server segments text according to the position of multimedia resources in the to-be-read information, converts the text segments into corresponding audio segments, and sequentially organizes the audio segments and the multimedia resources to generate a final reading file. The playing terminal can access the server and sequentially play the audio segments, thereby realizing voice reading service for user-defined information to be read, and timely displaying multimedia resources associated with the text in the process of text voice reading, which enhances the user's understanding of information. The above method quickly acquires and consumes various information through the playing terminal in combination with voice and partial vision without affecting the main task of the user, which not only improves the convenience of information acquisition, but also reduces security risks. BRIEF DESCRIPTION OF DRAWINGS
[0055] Figure 1A FIG. 1 is a schematic diagram of a three-party interaction scenario related to a cross-device information reading method according to an example embodiment of the present application;
[0056] Figure 1B is a step flow chart of a cross-device information broadcasting method applied to a server according to an example embodiment of the present application;
[0057] Figure 2 is an example of a text segmentation process according to an example embodiment of the present application;
[0058] Figure 3 is an example of a related sentence determination process according to an example embodiment of the present application;
[0059] Figure 4 is an example of a process of sequentially organizing audio segments and multimedia resources according to an example embodiment of the present application;
[0060] Figure 5A is a step schematic diagram of a content type determination of an indication message according to an example embodiment of the present application;
[0061] Figure 5B is an example of an indication message sending case in which a message field represents a link according to an example embodiment of the present application;
[0062] Figure 5C is an example of an indication message sending case in which a message field represents text according to an example embodiment of the present application;
[0063] Figure 6 is a step flow chart of a cross-device information broadcasting method for multi-party interaction according to an example embodiment of the present application;
[0064] Figure 7 is a structural schematic diagram of a cross-device information broadcasting device applied to a server according to an example embodiment of the present application;
[0065] Figure 8 is a hardware schematic diagram of an electronic device according to an example embodiment of the present application. DETAILED DESCRIPTION
[0066] The example embodiments will be described in detail herein with reference to the attached drawings. It should be understood that the terms first, second, third, etc. can be used to describe various information, but these information should not be limited to these terms. These terms are only used to distinguish one type of information from another type of information.
[0067] In the booming development of mobile Internet today, the user end such as a mobile phone becomes an important tool for obtaining life information. Users can obtain various information through mobile phone browsing web pages, reading articles, etc. at any time and any place. However, people's life scenes are complex and diverse, and are frequently switched. For example, in the driving scene, the traditional visual reading method such as looking at the mobile phone screen to obtain information is limited due to the need to focus on the road conditions. According to statistics, using a mobile phone to read information during driving can greatly increase the risk of traffic accidents. In some other scenes, such as when people are exercising (running, exercising, etc.), doing housework, or their eyes are tired and cannot visually read, a suitable information obtaining method is needed to meet the needs.
[0068] Current information obtaining technology mainly focuses on providing good user reading experience in static, visual reading suitable scenes. For example, most information applications mainly focus on how to optimize text layout, load pictures, etc. on the mobile phone screen to facilitate visual reading of users.
[0069] Therefore, there is an urgent need for an information obtaining method that can meet the dual needs of users for information obtaining convenience and safety, meet the continuous obtaining of mobile phone news information, etc. without interruption during driving, so as to obtain and consume various information such as news information, articles, etc. without affecting the main task of users.
[0070] The present application provides a cross-device information broadcasting method, which can provide cross-device voice broadcasting services for various information suitable for broadcasting services such as social media published article information, news reports, social dynamic posts, forum sharing, and public information such as public service notifications, as shown in Figure 1A The method involves the interaction of the user end, the service end, and the playing end.
[0071] The user end refers to a software or program client used by the user to forward or customize information that needs to be broadcasted. The software or program client runs on a smart phone, a tablet computer, a smart watch, a notebook computer, or other network-connected electronic devices. The electronic devices can be equipped with touch screens, keyboards, voice inputs, and other interactive interfaces to facilitate user input of information that needs to be broadcasted. The information input by the user end will be transmitted to the service end. For example, in the scenario of transmitting messages through WeChat platform services (such as public service numbers), the user end refers to the client (such as the input interface of the public service number) of the WeChat platform service facing users.
[0072] The service end refers to a device or platform for processing information sent by the user end and generating a broadcast reading file for the playing end to read and play, for example, it can be a cloud server, a data center or a dedicated broadcast reading service background. In the scenario of transmitting messages through the WeChat platform service (such as a public service number), the service end can be a custom server configured for the WeChat platform service, which can obtain an instruction message sent by the client from the WeChat platform service. The service end can establish a connection with multiple playing ends to facilitate users to read the broadcast reading file flexibly.
[0073] The playing end refers to a device for actually playing the information required to be broadcast and read, such as a television, a vehicle end such as a vehicle-mounted entertainment system, a smart speaker, a smart home device such as a smart mirror, a smart refrigerator with a display screen, or other electronic devices with audio / video playback functions. The playing end can access the service end to obtain the broadcast reading file and play it according to the characteristics of the device itself, such as screen size, speaker configuration, etc.
[0074] Through multi-party interaction, the broadcast reading file of news information, forum content, etc. is realized to flow between different devices, providing users with more convenient and diversified information acquisition methods, breaking the traditional mode of visual reading or page reading limited to the user end, so that users can easily transfer information to different devices for broadcast reading according to actual needs. For example, when a user finds an interesting news information on a smart phone outdoors but cannot read it, he / she can send it to the service end to process it into a broadcast reading file, and read the broadcast reading file through the vehicle end and play it by voice during the driving back home, meeting the user's needs in different scenarios.
[0075] Based on the above interaction scenario, the cross-device information broadcast reading method provided by the present application is described from the perspective of the service end, which is used to determine the information to be broadcast and read according to the received instruction message, and process it into a broadcast reading file for the playing end to play. For example, Figure 1B The method step flowchart is exemplarily shown. The implementation process of the cross-device information broadcast reading method at the service end can at least include the following steps:
[0076] S101, receiving an instruction message from a user end, and determining information to be broadcast and read according to the instruction message; the instruction message includes multi-modal information for positioning the information to be broadcast and read;
[0077] The indication message is a data packet sent by the user terminal to the server, which belongs to an operation instruction and is used to point to information or content that the user expects to read. The indication message is encapsulated according to the message transmission format defined between the user terminal and the server and then transmitted from the user terminal to the server. The indication message can be a web link of information that needs to be read, for example, when the user browses news or dynamics, the user can send the interested content to the input interface of the indication message through the sharing function provided by the software. Alternatively, the indication message can be the content itself of information that needs to be read, for example, when the user browses news or dynamics, the user can also copy the interested content and paste it into the input interface of the indication message.
[0078] The multi-modal information refers to information transmitted through different information representations or interactive modes such as text, image, sound, etc. The multi-modal information in the embodiment includes different types of information representations contained in the indication message, which can guide the server to locate the information to be read, that is, the user can generate the indication message through various modes such as voice instruction, picture, text input, etc. to accurately guide the server to find the information to be read. The input of multi-modal information can provide convenience for the user to input the indication message and reduce the use cost of the user.
[0079] The information to be read is used to represent the file that needs to be read and pointed to by the indication message. The information to be read can include text and multimedia resources such as pictures, audio or video obtained by analyzing the indication message. The information to be read can be text and multimedia resources extracted from the web link provided by the user, or information directly sent to the server by the user through the indication message.
[0080] Regarding the sending of the indication message from the user terminal to the server, the user terminal can send the indication message to the server through a self-developed application program, or can send the indication message through an existing social platform function, for example, a custom server can be implemented by using the public number interface provided by WeChat.
[0081] That is, the developer can develop a special application program according to specific requirements, which is connected to the server. Through the establishment of connection and data interaction mechanism, the application program can stably and efficiently send the indication message from the user terminal to the server. The function of sending the indication message to the server can be customized and developed according to specific business logic and functional requirements, to ensure that the sending of the indication message can accurately meet the requirements of the server for information acquisition. The user inputs related content in the application program interface, and the application program converts these inputs into the format of the indication message and sends them to the server according to the established communication protocol.
[0082] Alternatively, based on the social media, dynamic post, news information, etc. usually provide sharing to other software such as WeChat, and WeChat service platform such as WeChat public number or applet provides a series of interfaces for developers to realize the interaction with WeChat server, including the developers can customize the server to realize the deep integration with WeChat public number or applet. That is, the developer can create a customized WeChat public number or applet, and call the interface provided by WeChat to configure the custom server. When browsing social media or reading news, users can share interesting articles, posts or other content through the WeChat sharing function to the customized public number or applet, and the shared information is the indication information, which will be sent to the WeChat server through the input interface of the public number or applet, and forwarded to the custom server configured for the public number or applet by the WeChat server, so that the custom server processes the indication message. Through this way, the existing social media ecosystem can be used to reduce the operating cost, broaden the channel of user-end and server-end interaction, and provide more diversified choices for users.
[0083] When the server determines the to-be-read information according to the indication message, the content type of the indication message can be determined first, which is used to represent the way of indication message passing information, that is, the form in which the user specifies the to-be-read information, which can include any one of a web link and information text. Then, the indication message can be parsed according to the content type to determine the to-be-read information, and different parsing methods are used for different content types. For example, for a web link, the corresponding web content can be downloaded by accessing the URL (Uniform Resource Locator) in the indication message and the content is extracted to obtain the to-be-read information, and for information text, the information carried in the indication message can be obtained as the to-be-read information.
[0084] S102, according to the position of the multimedia resource in the to-be-read information, the text in the to-be-read information is segmented to obtain a plurality of text segments;
[0085] The text segmentation is used to divide the text part in the to-be-read information into smaller and independent text blocks according to the set logic. In this embodiment, the multimedia resource is used as the basis for text segmentation, and the multimedia resource is used as the text segmentation boundary. According to the position of the multimedia resource relative to the text in the to-be-read information, when the multimedia resource is detected, the text before and after the multimedia resource is divided into one or more independent text segments.
[0086] For example, assuming the information to be read out is “Hello everyone, let's take a look at the latest industry trends… Picture A… which shows the latest technology products developed by Company A… Video M… Today's sharing ends here, thank you for your listening!”, the information to be read out can be divided into at least three text segments according to Picture A and Video M, including text segment 1: “Hello everyone, let's take a look at the latest industry trends…”, text segment 2: “… which shows the latest technology products developed by Company A…”, and text segment 3: “… Today's sharing ends here, thank you for your listening.”
[0087] S103, converting the text segments into corresponding audio segments, and sequentially organizing the audio segments and the multimedia resources according to the position order of the text segments and the multimedia resources in the information to be read out to obtain a reading file, so that the playback end accesses the server to obtain the reading file, and the multimedia resources in the reading file are displayed in the process of playing the obtained reading file.
[0088] Regarding the conversion of the text segments into corresponding audio segments, text-to-speech technology, neural network text-to-speech technology, or other audio synthesis technology can be used to convert the text in each text segment into audio for voice reading. In the process of generating audio segments, a plurality of tone templates can be provided for the user, and the user can select the preferred tone template, so as to synthesize audio according to the tone characteristics of the tone template and the text of the text segment.
[0089] The above sequential organization of the audio segments and the multimedia resources is based on the relative position relationship between the text segments and the multimedia resources, that is, in the case of maintaining the information order of the text and the multimedia resources in the information to be read out, the corresponding audio segments of the plurality of text segments divided are sequentially integrated, and the multimedia resources are inserted into the corresponding audio segments of the text segments according to their relative positions, to obtain ordered data information as a reading file.
[0090] For example, for the above example of text segments 1-3, corresponding audio segments 1-3 can be obtained by speech-to-text technology, and the audio segments and the multimedia resources are sequentially organized according to the position order of the text segments and the multimedia resources in the information to be read out, that is, [text segment 1, picture A, text segment 2, video M, text segment 3], and the obtained playing content can be exemplarily represented as [audio segment 1, picture A, audio segment 2, video M, audio segment 3].
[0091] After obtaining the reading file corresponding to the indication message, the server can store the reading file in a preset reading storage queue or other storage structure or storage space, or the server can send the reading file to each playback end, so as to facilitate the user to select and play in the playback end.
[0092] For the case that the server receives multiple indication messages, the server can store each reading file as a reading sequence in the reverse order of receiving the indication messages after obtaining the reading file corresponding to each indication message, so that the playing end can obtain the reading file corresponding to the information that needs to be read and played recently when accessing the reading sequence, which facilitates the user to receive the latest information.
[0093] The playing end can access the server through a network request such as an HTTP request to request the reading file stored by the server. After receiving the request, the server can first verify the legality of the request, and send the reading file that has not been sent to the playing end for reading to the playing end after the legality verification is passed.
[0094] After receiving the reading file, the playing end can play the audio segments in the reading file in order from the first audio segment to the last audio segment based on the order of the audio segments in the playing content. In addition, the playing end can display multimedia resources at appropriate times during the playing of the audio segments to enhance user understanding. For example, based on the order of the audio segments and multimedia resources in the playing content, after the current audio segment is played, if the next adjacent content of the current audio segment is a multimedia resource such as a picture, the picture can be displayed and the audio segment after the multimedia resource can be played; if the next adjacent content of the current audio segment is a multimedia resource such as a video or audio, the video or audio can be played, and after the playing of the video or audio is completed, the audio segment after the multimedia resource can be played.
[0095] For example, the playing content can be exemplarily represented as [audio segment 1, picture A, audio segment 2, video M, audio segment 3], the playing end can play audio segment 1, audio segment 2 and audio segment 3 in order, and can also display picture A on the display device such as a display screen of the playing end when audio segment 1 is played, and continue to play audio segment 2, and display video M when audio segment 2 is played (the playing of audio segment is paused at this time), and continue to play audio segment 3 after the playing of video M is completed.
[0096] In the embodiments of the present disclosure, the user can send an indication message to the server through the user terminal at any time according to personal preferences and needs, customize a personalized broadcast file, and improve the satisfaction of user experience and the flexibility of information acquisition. The server can segment the text according to the multimedia resource position in the to-be-broadcast information, convert the segmented text into audio segments, organize the audio segments and multimedia resources in order to obtain a broadcast file, so that the information can be received by the user in an auditory manner. The audio segments are sequentially played through the playback terminal, and the multimedia resources are displayed, cross-platform and multi-device information flow are realized, visual reading is not required, the distraction of the user due to reading is reduced, the user can safely acquire and consume information without affecting the main task, thereby adapting to the needs of the user in different life scenarios, providing a smooth and coherent information broadcast experience, and enhancing the safety of information consumption.
[0097] The method can meet the needs of users in various scenarios (including but not limited to driving scenarios), so that the user can completely and continuously acquire information even in the case of being unable to visually read, thereby improving the safety and efficiency of information acquisition. For scenarios where attention or both hands are limited or other visual reading is inconvenient, such as during driving, exercising, or busy with housework, the user can send to-be-broadcast information or news to the server through the mobile terminal for processing into a broadcast file, and play the broadcast file output by the server through the playback terminal such as the car machine terminal during driving to continue the user's "reading", which is not affected by the physical environment or time limit, thereby enhancing the flexibility and continuity of user experience.
[0098] For example, in a driving scenario, the car owner can transfer information on the mobile phone to the vehicle system through the method, and continue to acquire information through a combination of voice listening and multimedia resource display on the car machine terminal. In other scenarios, such as a sports scenario, the user can acquire information without interrupting the sports, thereby improving the accessibility of information and user experience.
[0099] In some embodiments, since the WeChat platform has a wide user base, it provides a perfect interface that allows developers to customize servers for public service numbers or small programs, and supports multi-modal input such as text, voice, and pictures, which can well receive more forms of information content. Correspondingly, the mobile terminal information such as articles, news, and posts also supports sharing information in the application to the WeChat platform. Therefore, in order to improve the convenience and immediacy of communication between the user terminal and the playback terminal, the present embodiment takes the WeChat platform service as the carrier of the indication message reception, and takes the custom server configured by the developer on the WeChat platform service as the server for processing the indication message to obtain the to-be-broadcast information, wherein the WeChat platform service can include but is not limited to public service numbers and small program applications.
[0100] Based on this, the service end receiving the indication message from the user end in the foregoing step S101, that is, the service end receives the indication message forwarded by the WeChat server; the WeChat server receives the indication message sent by the user through the client of the WeChat platform service. Wherein, the WeChat server as a transfer between the WeChat platform service and the custom server configured by the developer, forwards the indication message from the input interface of the WeChat platform service to the custom server.
[0101] That is, the user can send the indication message through the input / chat interface of the client of the WeChat platform service such as WeChat public number, applet, etc., such as directly sharing the information to be broadcasted to the designated WeChat public number of WeChat through the function of sharing to WeChat provided by the information software, or copying the information content and sending after copying in the input box of the input / chat interface, etc., to transmit the content to be broadcasted.
[0102] After the WeChat server receives the indication message sent by the client of the WeChat platform service, it forwards it to the custom server configured by the developer on the WeChat platform service according to the pre-agreed protocol or rule, so that the custom server receives the indication message forwarded by the WeChat server, determines the information to be broadcasted according to the indication message, and further generates the broadcast file.
[0103] In the embodiments of the present disclosure, based on the WeChat platform, a perfect interface and transmission mechanism are provided, through the WeChat platform service as a transmission medium, the user can directly interact with the custom server through WeChat, without the need to download additional third-party applications for communication with the playback end, simplifying the user's use process, reducing the user's use cost, avoiding the high cost brought by developing and maintaining independent communication applications. At the same time, the WeChat server as a transfer and the WeChat platform with instant messaging capability, forwards the indication message from the input interface of the WeChat platform service to the custom server, without additional transmission cost or complex transmission protocol, thereby reducing the transmission cost and ensuring that the indication message can be efficiently and accurately transmitted from the user end to the custom server.
[0104] In some embodiments, the server segments the text in the to-be-read information to obtain a plurality of text segments, as described in step S102 of the foregoing embodiments. This embodiment provides a way of detecting multimedia resources to segment text during traversal of the to-be-read information, that is, the to-be-read information is traversed, and when text is detected, the text is continuously collected until a multimedia resource is detected or the traversal ends, and the currently collected text is taken as a text segment. That is, the to-be-read file is detected from the starting position, and when pure text content is detected, the text characters are continuously collected. The collection process continues until a multimedia resource is detected or the traversal of the to-be-read information ends, at which time the continuously collected text is divided into a text segment. Meanwhile, when a multimedia resource is detected, the multimedia resource can also be associated with the recently collected text segment, which is the text segment obtained by continuous collection when the multimedia resource is detected. The association can be based on a timestamp, a sequential number, or any way that can ensure accurate correspondence during playback. The association between the audio segment and the multimedia resource can be achieved by recording the correspondence between the audio segment and the multimedia resource in an internal data structure such as a hash table or a linked list.
[0105] For example, as Figure 2 An example of a text segmentation process is shown. Taking the to-be-read information "Hello everyone, let's take a look at the latest industry trends… Picture A… It shows the latest technology products developed by Company A… Video M… Today's sharing ends here, thank you for your listening!" as an example, the to-be-read information is detected from the starting position, and the text is continuously collected until Picture A is detected, and the collected "Hello everyone, let's take a look at the latest industry trends…" is taken as text segment 1, and Picture A is associated with the recently collected text segment 1. Based on the same principle, the collected "… It shows the latest technology products developed by Company A…" is taken as text segment 2, and Video M is associated with text segment 2, and then "… Today's sharing ends here, thank you for your listening." is taken as text segment 3, and the traversal of the to-be-read information ends.
[0106] Based on the above-mentioned way of associating the multimedia resource with the recently collected text segment, for the server to sequentially organize the audio segments and the multimedia resources to obtain the read file as described in step S102 of the foregoing embodiments, the associated multimedia resource can be added at the end of the playback of each audio segment, and the read file is integrated.
[0107] In this embodiment, a data structure about the read-aloud file can be created, which can accommodate a sequence of audio segments and multimedia resources, and can be a list or a set or other applicable structure, and the elements include audio segments or markers representing multimedia resource playback points. By traversing the association records of the text segments and multimedia resources, for each audio segment generated from a text segment, it is added to the data structure of the read-aloud file, and according to the association records, a marker representing the multimedia resource playback point is inserted in the data structure after each audio segment, which can contain the reference of the multimedia resource (such as URL or file path) and / or any additional information required to play the resource (such as play duration, format, etc.). Based on the added markers, when playing to the marker representing the multimedia resource playback point at the playback end, the audio playback can be paused, and the corresponding multimedia resource can be loaded and displayed.
[0108] In the process of adding the associated multimedia resources to the end of the playback of each audio segment and integrating the read-aloud file, the order of the audio segments and multimedia resources is consistent with the original order in the information to be read, so as to obtain coherent information. The server generates a read-aloud file containing audio segments and multimedia resource playback points, and the playback end can accurately load and display the multimedia resource when playing to the corresponding position, thereby providing a rich information receiving experience combining hearing and vision for the user.
[0109] In the embodiments of the present disclosure, the text and multimedia resources such as pictures, videos or audio clips contained in the information to be read are detected simultaneously in the process of traversing the information to be read, and the multimedia resources are accurately associated with the text segments, so that in the process of playing the audio segments, the related multimedia resources can be displayed synchronously at the appropriate time of playing the corresponding audio segments, thereby enhancing the hearing and vision experience of the user.
[0110] In some embodiments, in order to better improve the user experience and the coherence of the content, the present embodiment not only focuses on the direct association of audio segments and multimedia resources, but also further identifies the inherent relationship between multimedia resources and adjacent text content. The present embodiment provides another way of associating multimedia resources with audio segments, which ensures that the user can naturally transition to the related multimedia resources when listening to the audio content, while maintaining the logicality and integrity of the information.
[0111] In this embodiment, after obtaining a plurality of text segments, the server determines at least one text segment adjacent to each multimedia resource for the multimedia resource; next, the graphic-text relevance of each sentence in the multimedia resource and the adjacent text segments is obtained, and the relevant sentence with the highest graphic-text relevance is determined from each sentence.
[0112] The at least one text segment adjacent to the multimedia resource can include two text segments obtained by taking the multimedia resource as a division basis. The image-text correlation degree can be used to measure the relevance or matching degree between the multimedia resource and each of the sub-sentences in the adjacent text segments, and can be measured based on the similarity of content, the consistency of semantics, and the degree of fit of context or theme.
[0113] As Figure 3 An example of a relevant sub-sentence determination process is shown. The multimedia resource is adjacent to text segments P and Q. For each adjacent text segment, taking text segment P as an example, the text segment can be divided into multiple sub-sentences (e.g., X sub-sentences in the figure) in units of sentences. The server can use image-text algorithms, feature matching, semantic matching, machine learning, or other feasible methods to deeply analyze the content of the sub-sentences in the text segment, such as keyword extraction, sentiment analysis, and theme recognition. At the same time, the features of the multimedia resource, such as image content and label information, are combined to quantify the relevance between the two, calculate the image-text correlation degree between the multimedia resource and each of the sub-sentences in each adjacent text segment, and obtain X image-text correlation degrees. By accurately calculating the image-text correlation degree, the server can select the relevant sub-sentence with the highest image-text correlation degree with the multimedia resource from the X image-text correlation degrees. The relevant sub-sentence best represents the core content of the multimedia resource or forms a complement with the multimedia resource, thereby ensuring that the playback end inserts the multimedia resource when the audio is played to the text segment most relevant to the multimedia resource, thereby providing the user with a more coherent and lively reading experience. Assuming that the second image-text correlation degree in the X image-text correlation degrees represents the highest degree of relevance between the multimedia resource and the second sub-sentence, the second sub-sentence is selected as the relevant sub-sentence.
[0114] Based on this, the server sequentially organizes the audio segments and the multimedia resource to obtain the reading file in step S102 of the foregoing embodiment. The relevant sub-sentence corresponding to the multimedia resource can be implemented, i.e., the audio playback position of the relevant sub-sentence can be determined for the audio segment corresponding to the adjacent text segment containing the relevant sub-sentence, and the multimedia resource can be added at the audio playback position of the relevant sub-sentence corresponding to the multimedia resource, and the reading file can be integrated.
[0115] For adding the multimedia resource at the audio playback position, the audio segment can be imported into an editing timeline through a video or audio editing application, and the edited file after inserting the multimedia resource at the audio playback position can be exported as the audio segment after adding the multimedia resource. Alternatively, a playback script can be created to specify the playback order and time points of the audio segment and the multimedia resource. In addition, any multimedia resource adding method that can dynamically insert and display the multimedia resource during audio playback can be used, and the present application does not limit the method.
[0116] In the embodiments of the present disclosure, by considering the context relationship and semantic correlation of the text content, the multimedia resource is inserted at the audio playing position of the relevant sub-sentence corresponding to the multimedia resource, so that the selected audio segment can accurately reflect the background or supplementary information of the multimedia resource, the accurate integration of the audio and the multimedia resource is realized, the information conveying effect and the depth of user experience are enhanced, and the coherence of the audio stream and the fluency of the user experience are maintained.
[0117] In the foregoing embodiments, how to add multimedia resources at the specified playing position of the audio segment is described to enhance the expressiveness and richness of the reading file, and facilitate the reading file to more accurately reflect the structure and intention of the original information. In actual application, as the complexity of the reading file increases, such as in the case of containing a large number of text segments and multimedia resources, a systematic and efficient method is needed to manage and organize these elements so that the finally generated reading file is both coherent and orderly.
[0118] To this end, the embodiments provide a data slice management method based on unique identification, which is used to more flexibly and efficiently organize audio segments and multimedia resources, simplifies the data processing process, and improves the scalability and maintenance flexibility of the reading file. Specifically, the process of the server sequentially organizing audio segments and multimedia resources to obtain a reading file can be implemented in the following manner: setting a unique identification for each text segment; organizing the audio segment corresponding to the text segment, the added multimedia resource and its resource type into a data slice with the unique identification; and sequentially integrating all data slices according to the position order of the text segments to form a complete reading file.
[0119] The unique identification is used to distinguish different text segments and can be an incremental integer or any other form of unique string, which can serve as a unique index of the text segment or the audio segment.
[0120] The server encapsulates the audio segment corresponding to each text segment, the multimedia resource added in the audio segment and its resource type (such as pictures, videos, etc.) into an independent data slice together with the unique identification. Each data slice belongs to a small data packet containing all related information, stores specific media content, and records the association relationship and playing order between them.
[0121] For example, one of the exemplary data slice structures is as follows, including a field sort representing the unique identification, a field content representing the text segment content, a field mediaType representing the multimedia resource type, a field assetsSrc representing the multimedia resource, and a field audio representing the audio segment:
[0122] {“sort”: 1,
[0123] “content”: text segment 1,
[0124] “mediaType”: resource type such as picture or video or audio,
[0125] “assetsSrc”: multimedia resource itself or link or storage address,
[0126] “Audio”: audio segment 1 converted from text segment 1 or link or storage address.
[0127] After each audio segment is organized into a corresponding data slice, the data slices are integrated one by one in the natural order of the original text segment in the to-be-read information to form a coherent read file stream, maintaining the logicality and coherence of the content and the seamless connection of multimedia resources and audio content, making the entire read experience smooth and rich.
[0128] In the process of organizing the read file in the above manner, if it is detected that a certain audio segment has defects or there is a need to add a new audio segment, the data slice that needs to be modified or updated can be quickly located according to the unique identifier, and the field information recorded in the data slice is replaced or supplemented accordingly, thereby realizing the quick update of the read file and improving the scalability and maintenance flexibility of the read file. For example, if the audio segment has a problem, the field value representing the audio segment resource can be replaced with a new, defect-free audio segment; if a new audio segment needs to be added, a data slice about the new audio segment can be organized and added to the read file, and the logical coherence with adjacent content is ensured.
[0129] In the embodiments of the present disclosure, by introducing the concepts of unique identifier and data slice, the organization mode of audio segment and multimedia resource is optimized, and the flexibility and reliability of cross-device information read service are improved, which is suitable for processing content containing a large amount of text and multimedia elements, ensuring the accuracy and consistency of the read result, and providing users with a more smooth and immersive read service experience.
[0130] In some embodiments, for a user in a specific environment or situation, such as focusing on driving, sports, etc., cannot directly view picture or video content, or there are safety hazards in the scene, in order to meet the needs of the user in a specific scene, while ensuring the effectiveness and convenience of information transmission, the present embodiment innovatively provides a new way of sequentially organizing audio segments and multimedia resources, conveying information that is originally presented visually through audio form.
[0131] That is, in the process of sequentially organizing audio segments and multimedia resources to obtain the reading file, in the case that the multimedia resource is a picture, the server in this embodiment can generate a description text of the picture through a picture-text network, convert the description text into audio to obtain an audio segment corresponding to the description text, and then sequentially organize the audio segments corresponding to the text segments and the multimedia resource to obtain the reading file according to the position of the multimedia resource relative to the text segments, wherein the picture-text network can deeply analyze and understand the picture and automatically generate a description text accurately describing the content of the picture, so that the user can accurately receive the key information conveyed by the picture according to the description text even in the case that the user cannot directly view the picture.
[0132] In the case that the multimedia resource is a video, an audio file in the video can be extracted as an audio segment, and the audio segments corresponding to the text segments and the multimedia resource can be sequentially organized to obtain the reading file, wherein the audio file in the video can include all sound elements in the video such as dialogues, narrations, background music and the like, and the sound elements collectively constitute an audio segment corresponding to the video; or the video can be processed through a picture-text network to generate a description text or a summary that can summarize the key content of the video, and the description text corresponding to the video can be converted into an audio format as another audio segment corresponding to the video, so that in the reading file, the sound information of the video can be received through hearing, and the main content of the video can be understood through the audio form of the description text, thereby further improving the richness of the reading file and the transmission efficiency of information.
[0133] For example, as Figure 4 An example of a process of sequentially organizing audio segments and multimedia resources is exemplarily shown. Taking the position sequence of the text segments and the multimedia resources of the aforementioned to-be-read information [text segment 1, picture A, text segment 2, video M, text segment 3] as an example, a descriptive text of the picture A can be generated through a picture-text network, which is converted into an audio segment A, the original audio file of the video M can be extracted as an audio segment M1, and a key information summary of the video M can be generated through a picture-text network, which is converted into an audio segment M2, then after obtaining the audio segments 1-3 corresponding to the respective text segments, the sequential organization of the audio can be exemplified as [audio segment 1, audio segment A, audio segment 2, audio segment M1, audio segment M2, audio segment 3].
[0134] The foregoing embodiment describes that the server can determine the content type of the indication message first when determining the to-be-read information according to the indication message, so as to parse the indication message according to the content type to obtain the to-be-read information. Regarding the determination process of the content type of the indication message, this embodiment proposes a manner of determining the content type based on the field value in the transmission packaging format of the indication message, such as Figure 5AThe content type determination step shown in the schematic diagram, determining the content type of the indication message can be realized by the following way:
[0135] S501, obtaining the field value of the message type in the transmission encapsulation format of the indication message;
[0136] Transmission encapsulation format is used to describe the format adopted by data when transmitting in the communication network, including the structure, encoding method, frame format of data, etc. The encapsulation format ensures that the data can be correctly transmitted in the network and can be correctly parsed at the receiving end. Different network protocols or systems may adopt different encapsulation formats.
[0137] The field is a unit in the data structure, used to store data of a specific type. In the communication protocol, the message is usually organized into multiple fields, each field contains different types of information such as message type, source address, target address, etc. The length and position of the field are usually defined by the protocol specification.
[0138] The data transmission encapsulation format is different under different communication protocols, and the corresponding field representing the message type will also change. In actual application process, it can be determined according to the communication protocol adopted.
[0139] Taking the user sending news information to the custom WeChat public number as an example, the WeChat public number client forwards to the WeChat server, and then to the service end of the application for processing indication message, the WeChat server submits the XML data packet of the client sent message to the service end of the application. The transmission encapsulation format of the indication message can be represented as follows:
[0140]
[0141]
[0142] Among them, the field <msgtype>< / msgtype> is used to represent the message type, and the field value of the message type in the XML data packet of this example is text, that is, the user sends the indication message by inputting text in the input box of the public number.
[0143] S502, in the case that the field value represents a link, determining that the content type is a web link;
[0144] Regarding the case that the content type of the indication message is a link, such as the user in the news information, dynamic post, etc. The content page that wants to be broadcasted and read shares it in the form of web technology link to the sending interface of the indication message through the sharing function of other software, and then it is automatically transmitted to the service end. The corresponding field value of the message type will point to the link, such as the field value of link.
[0145] Still taking the example that the WeChat server submits the XML data packet of the message sent by the client to the server of the application, the indication message is sent in the form of a link as shown in Figure 5B The XML data packet includes <msgtype><![CDATA[ link ]]> < / msgtype> , the field value of the message type is link, that is, the field value indicates a link, and it is determined that the content type of the indication message is a web link.
[0146] S503, in the case where the field value indicates text, whether the string in the indication message includes a link in the form of text is detected by a regular expression:
[0147] S5031, if yes, it is determined that the content type is a web link;
[0148] S5032, if no, it is determined that the content type is information text.
[0149] The field value indicates text, which means that the user inputs information through the input box on the sending interface of the indication message and sends it to the server as an indication message. Based on the fact that the user can copy the link address of the content to be read or copy the content itself and paste the copied content in the input box as an indication message, in the case where the field value indicates text, it is necessary to further distinguish whether the text itself belongs to a link address or a file to be read.
[0150] The embodiment detects whether the text characters carried by the indication message include a link by a regular expression that detects strings conforming to the URL format. Based on the fact that a URL at least includes a protocol (such as http: / / , https: / / , ftp: / / , etc.), a domain name, a path, a query parameter, and the like, an exemplary regular expression can be set as https?:\ / \ / [^\s]+, which is used to match a URL starting with http or https. Among them, “https?” is used to match http or https, and “?” indicates that the character “s” can appear 0 times or 1 time; “:\ / \ / ” is used to match: / / , which is a part that must be followed after the separator between the protocol part and the domain name part in the URL protocol; and “[^\s]+” indicates matching one or more continuous non-space characters (\s represents a space character, including a space, a tab, a line feed, etc.).
[0151] S504, in the case where the field value indicates an image, text information is extracted from the image, and whether the extracted text information includes a link in the form of text is detected by a regular expression:
[0152] S5041, if yes, it is determined that the content type is a web link;
[0153] S5042, if not, then determine that the content type is information text.
[0154] Extracting text information from an image can be achieved using OCR (Optical Character Recognition) technology. Before extracting the text, to improve the accuracy of OCR, the image can be preprocessed, such as by converting it to grayscale, binarizing it, or removing noise. After extracting the text information, the regular expressions described in the previous steps can be used to identify webpage links to determine whether the text information is a text-based link, thus accurately determining the content type of the indicated message.
[0155] like Figure 5C The field values shown in the instruction message represent examples of text transmission status. Using the aforementioned XML data packet as an example, the XML data packet will include... <msgtype><![CDATA[ text ]]> < / msgtype> Regular expressions are used to extract values from fields that represent the text content carried in the instruction message. <content>< / content> Link detection is performed. If the transmitted instruction message content is "https:mp.weiixn.qq.com / s / tv_VuNmhqQs_r747mlqeGA", and a regular expression match is successful, then the content type is determined to be a webpage link. If the transmitted instruction message content is "At 3:52 AM today, something happened at the wonton shop on the first floor of Yongxin Building...", and a regular expression match fails, then the content type is determined to be information text, which may include text and multimedia resources.
[0156] In this embodiment of the disclosure, by obtaining the value of the field representing the message type in the transmission encapsulation format of the indication message, it is possible to initially and quickly determine whether the user is sending a link or text information. Then, if the field value represents text, further link detection is achieved through regular expressions, effectively distinguishing whether the input is a text link or the content to be played, thereby significantly improving the accuracy of content type identification of the indication message.
[0157] Different parsing methods will be used for different content types of the instruction message to accurately obtain the content that needs to be played. This embodiment provides the parsing methods for content types of information text and web page links, namely:
[0158] When the content type is information text, the step of parsing the instruction message according to the content type to determine the information to be played can be done by obtaining the value of the field representing the message content in the transmission encapsulation format of the instruction message as the information to be played.
[0159] In the case that the content type is a webpage link, the webpage content can be located according to the link in the indication message, i.e. the value of a field representing the message content in the transmission encapsulation format of the indication message is used to locate the webpage content; next, the webpage structure of the webpage content is parsed, and the webpage text and multimedia resources are sequentially obtained as the to-be-read information according to the relative position relationship between the webpage text and the multimedia resources.
[0160] In the case that the content type is a webpage link, the webpage content can be located according to the link in the indication message, i.e. the value of a field representing the message content in the transmission encapsulation format of the indication message is used to locate the webpage content; next, the webpage structure of the webpage content is parsed, and the webpage text and multimedia resources are sequentially obtained as the to-be-read information according to the relative position relationship between the webpage text and the multimedia resources. <title>tag as article title, identify< / title> 、 、 <section>< / section> 、 、 <video>Tags such as tags, web page information extraction.
[0161] For example, assume that the HTML web page static code of a certain news web page is as follows:
[0162]
[0163] For the web page structure, each HTML tag can be sequentially identified, from <h1>The "news headline: New developments in artificial intelligence" is extracted from the tags, and the text and multimedia resources are extracted in order: ① "This is the first paragraph of the news." ② The image "AI image" in the tags; ③ "This is the second paragraph of the news, which mentions some research results about AI." ④ <video>the third paragraph of the news text, which summarizes the current development trend of AI. Among them, the position and resource link of the multimedia resource can be recorded. The extracted text and multimedia resources are combined in order to form an ordered information list as a to-be-read file. In this way, the structure of the original document and the position relationship of the multimedia elements can be preserved, facilitating subsequent text segmentation processing and sequential organization of audio segments and multimedia resources.
[0164] After describing the cross-device information reading method provided by the application from the perspective of the server, the cross-device information reading method will be further described from the perspective of the interaction between the user side, the server side and the playing side. For example, Figure 6 The method steps flowchart is exemplarily shown. The cross-device information reading method provided by the application can at least include the following steps:
[0165] S601, the user side sends an indication message to the server to make the server determine the to-be-read information according to the indication message; the indication message includes multi-modal information for positioning the to-be-read information;
[0166] In the cross-device information reading method provided by the application, the user side sends the user-defined or source from social media and the like to the sending interface of the indication message, and sends the to-be-read file to make the server determine the to-be-read information according to the content type of the received indication message.
[0167] For example, when a user browses a news webpage, the user can forward the news of interest to the sending interface of the indication message through the function of sharing to other software provided by the news webpage, and send it to the server as an indication message. Alternatively, the user can copy the webpage link of the news of interest or copy the news content itself, paste it in the input box of the sending interface of the indication message, and send it to the server as an indication message.
[0168] S602, the server segments the text in the to-be-read information according to the position of the multimedia resource in the to-be-read information to obtain a plurality of text segments, converts the text segments into corresponding audio segments, and sequentially organizes the audio segments and the multimedia resources according to the position order of the text segments and the multimedia resources in the to-be-read information to obtain a reading file;
[0169] S603, the playing side accesses the server to obtain the reading file, and displays the multimedia resources in the process of sequentially playing the audio segments.
[0170] In the cross-device information broadcast method provided in the present application, the playing end and the service end establish a communication connection, can actively access the service end to request a broadcast file, and store the requested broadcast file to the playing end local for playing. For any broadcast file, the playing end can display the added multimedia resource on the display device of the playing end when the audio segment is played to the playing position where the multimedia resource is added, in the process of sequentially playing each audio segment included in the broadcast file.
[0171] When the cross-device information broadcast method is applied to the playing end, the playing end can access the service end through an interface to obtain a broadcast file to be broadcast, and sequentially play the audio segments in the broadcast file, and synchronously display the multimedia resource related to the audio segment on the display device. In the process of synchronously displaying the multimedia resource related to the audio segment, the resource type of the multimedia resource added by the audio segment can be detected, in the case that the resource type represents a picture, a picture pop-up window can be displayed on the display device, and the pop-up window content is a specific picture, in the case that the resource type of the multimedia resource is an audio or video containing sound, the playing of the audio segment is stopped and a pop-up window is displayed to play the multimedia resource, until the multimedia resource is played to completion and the playing of the audio segment is continued.
[0172] The specific implementation process of steps S601-S603 can be referred to the description of the foregoing embodiments, which will not be repeated here.
[0173] In order for those skilled in the art to better understand the cross-device information broadcast method provided in the present application, the following user browses news information through the user end, and expects to broadcast the news information on the car machine end. The scene is taken as an example to exemplarily illustrate the cross-device information broadcast method.
[0174] The user can upload the content of the news information by sending an instruction message. In order to ensure the convenience and immediacy of the communication between the car owner's mobile phone and the car machine, a WeChat public service number can be used as a carrier for receiving news information, so that the car owner user does not need to download other apps, which reduces the user's use cost and the application research and development cost of the car enterprise. The user sends the web link of the news information to the WeChat public number for realizing cross-device information broadcast through the WeChat public chat window, so as to complete the uploading of the information. The WeChat server sends the received XML data packet of the instruction message to the service end URL customized by the developer, so that the instruction message can be sent to the service end, and the service end can be located to the news information to be broadcast according to the instruction message.
[0175] The service end can determine the message type of the received instruction message according to the field indicating the message type in the XML data. <msgtype>< / msgtype> value, and the regular expression accurately identifies the content type of the message. In the case of the content type being a link, the webpage final static code can be obtained through a crawler technology, and the news information content can be extracted as the to-be-read information by parsing the HTML tag. In the case of the content type being information text, the field <content>< / content> value is obtained, and the to-be-read information is taken as the to-be-read information.
[0176] Based on the fact that the to-be-read information includes media resources such as pictures and videos as content to assist readers in understanding the article content, the multimedia resources and the text together constitute the news information content, and therefore the media content such as pictures and videos can be displayed on the vehicle end at a suitable time when the text content of the news information is played. Therefore, in this embodiment, the server does not generate an audio file for the entire news information, but instead seamlessly connects the audio segment corresponding to the text segment and the multimedia resource by means of text segmentation + marking the media type and resource, so as to facilitate the display of the media content at the correct time when the audio segment is played. By sequentially organizing the audio segments and the multimedia resources, a reading file is obtained, and the elements in the reading file include audio and multimedia resources, which are used for reading on the vehicle end to realize the voice reading service of the news information.
[0177] After the vehicle end is started, the vehicle end can access the server through an interface API to read at least one reading file, and sequentially play each reading file in the order of the playing queue of the reading file in the server. For any reading file, the first audio segment is played in the order of the audio segments and the multimedia resources in the reading file. During the playing of the audio segment, if it is detected that the audio segment is added with a multimedia resource, the multimedia resource can be displayed through a pop-up window at the audio playing position where the multimedia resource is added, or after the playing of the audio segment is completed, until all the audio segments in the playing content are played.
[0178] This embodiment focuses on the driving scene, skillfully combines the convenience of mobile devices and the intelligence of voice technology, realizes the flow and seamless connection of news information on different device ends, allows the vehicle owner to easily transfer the information content in the mobile phone to the vehicle system during driving, and through the voice playing function, the news information can be continued in the fragmented time of driving, not only ensuring the driving safety, but also improving the accessibility and reading efficiency of information. Through this method, the vehicle owner is no longer limited to the traditional visual reading method, and can also enjoy the pleasure of reading through voice listening while holding the steering wheel and focusing on the road conditions, broaden the knowledge horizon, and enrich the driving journey.
[0179] Corresponding to the foregoing embodiment of the cross-device information reading method, the application provides a cross-device information reading system, which comprises a user terminal, a server and a playing terminal; wherein the user terminal sends an indication message to the server so that the server determines to-be-read information according to the indication message; the server processes text in the to-be-read information according to the position of multimedia resources in the to-be-read information to obtain a plurality of text segments, converts the text segments into corresponding audio segments, and sequentially organizes the audio segments and the multimedia resources according to the position sequence of the text segments and the multimedia resources in the to-be-read information to obtain a reading file; and the playing terminal accesses the server to obtain the reading file and displays the multimedia resources in the reading file in the process of playing the reading file, i.e. displays the multimedia resources in the reading file in the process of sequentially playing the audio segments in the reading file.
[0180] The application also provides another cross-device information reading device, which is applied to a server, as shown in the figure, and comprises: Figure 7
[0181] An information obtaining module 701 is configured to receive an indication message from a user terminal and determine to-be-read information according to the indication message; the indication message comprises multi-modal information for locating the to-be-read information;
[0182] A text segment processing module 702 is configured to process text in the to-be-read information according to the position of multimedia resources in the to-be-read information to obtain a plurality of text segments;
[0183] A reading file construction module 703 is configured to convert the text segments into corresponding audio segments, sequentially organize the audio segments and the multimedia resources according to the position sequence of the text segments and the multimedia resources in the to-be-read information to obtain a reading file, so that a playing terminal accesses the server to obtain the reading file and displays the multimedia resources in the process of sequentially playing the audio segments.
[0184] In some embodiments, the text segment processing module is specifically configured to:
[0185] traverse the to-be-read information, continuously collect text until a multimedia resource is detected or the traversal ends, and take the currently collected text as a text segment; and in the case that a multimedia resource is detected, associate the multimedia resource with the recently collected text segment;
[0186] The reading file construction module is specifically configured to add the associated multimedia resource at the end of the playing of each audio segment, and integrate to obtain the reading file.
[0187] In some embodiments, the apparatus further comprises: a server determining, for each multimedia resource, at least one text segment adjacent to the multimedia resource; obtaining the multimedia resource and the graphic-text relevance of each sentence in the adjacent text segment, and determining the most relevant sentence with the highest graphic-text relevance therefrom;
[0188] The broadcast file construction module is specifically configured to: determine the audio playing position of the relevant sentence for the audio segment corresponding to the adjacent text segment; add the multimedia resource at the audio playing position of the relevant sentence corresponding to the multimedia resource, and integrate to obtain the broadcast file.
[0189] In some embodiments, the broadcast file construction module is specifically configured to:
[0190] setting a unique identifier for each text segment; organizing the audio segment corresponding to the text segment, the added multimedia resource and the resource type thereof, and the unique identifier into a data slice; sequentially integrating all data slices according to the position order of the text segments to form a complete broadcast file.
[0191] In some embodiments, the broadcast file construction module is specifically configured to:
[0192] In the case where the multimedia resource is a picture, generating a description text of the picture through a graphic-text network, obtaining an audio segment corresponding to the description text, and sequentially organizing the text segment and the audio segment corresponding to the multimedia resource to obtain the broadcast file; or, in the case where the multimedia resource is a video, extracting an audio file in the video as an audio segment, and sequentially organizing the text segment and the audio segment corresponding to the multimedia resource to obtain the broadcast file.
[0193] In some embodiments, the information acquisition module is specifically configured to: determine the content type of the indication message; the content type includes any one of a webpage link and information text; parse the indication message according to the content type to determine the information to be broadcast.
[0194] In some embodiments, the information acquisition module comprises, when used to determine the content type of the indication message:
[0195] obtaining a field value representing a message type in a transmission encapsulation format of the indication message; in a case that the field value represents a link, determining that the content type is a web link; in a case that the field value represents text, detecting whether a string in the indication message includes a link in text form by using a regular expression, and if yes, determining that the content type is a web link, and if no, determining that the content type is information text; in a case that the field value represents an image, extracting text information from the image, and detecting whether the text information belongs to a web link or information text by using a regular expression.
[0196] In some embodiments, the information obtaining module, when used to determine the to-be-read information by parsing the indication message according to the content type, comprises:
[0197] in a case that the content type is a web link, locating to web content according to a link in the indication message; parsing a web structure of the web content, and sequentially obtaining web text and multimedia resources as the to-be-read information according to a relative position relationship between the web text and the multimedia resources;
[0198] or, in a case that the content type is information text, obtaining a field value representing message content in a transmission encapsulation format of the indication message as the to-be-read information.
[0199] In some embodiments, the server comprises a custom server on a WeChat platform service; and the information obtaining module, when used to receive the indication message from the user terminal, comprises: receiving the indication message forwarded by a WeChat server; and the WeChat server receives the indication message sent by the user through a client of the WeChat platform service.
[0200] The implementation processes of the functions and roles of the units in the above apparatus are specifically described in the implementation processes of the corresponding steps in the above method, and will not be described here.
[0201] Embodiments of the present application also provide an electronic device, a structure diagram of which is shown in Figure 8 The electronic device 800 comprises at least one processor 801, a memory 802 and a bus 803, the at least one processor 801 is electrically connected with the memory 802; the memory 802 is configured to store at least one computer executable instruction, and the processor 801 is configured to execute the at least one computer executable instruction, so as to execute the steps of any one cross-device information reading method provided by any one embodiment or any one optional implementation manner of the present application.
[0202] Further, the processor 801 can be an FPGA (Field-Programmable Gate Array) or other device with logic processing capability, such as an MCU (Microcontroller Unit) or CPU (Central Process Unit).
[0203] The embodiment of the present application further provides another readable storage medium storing a computer program, and the computer program is used to implement the steps of any cross-device information broadcasting method provided by any one of the embodiments or any one of the optional implementation manners.
[0204] The readable storage medium provided by the embodiment of the present application includes but is not limited to any type of disk (including a floppy disk, a hard disk, an optical disk, a CD-ROM, and a magneto-optical disk), a ROM (Read-Only Memory), a RAM (Random Access Memory), an EPROM (Erasable Programmable Read-Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a flash memory, a magnetic card or an optical card. That is, the readable storage medium includes any medium storing or transmitting information in a readable form by a device (for example, a computer).
[0205] The above merely provides the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.< / video> < / h1> < / video>
Claims
1. A cross-device information casting method, characterized by, Applied to a server side, the method comprises: receiving an indication message from a user side, and determining to-be-read information according to the indication message; the indication message comprises multi-modal information for positioning the to-be-read information; segmenting text in the to-be-read information according to positions of multimedia resources in the to-be-read information, to obtain a plurality of text segments; converting the text segments into corresponding audio segments, and sequentially organizing the audio segments and the multimedia resources according to a position sequence of the text segments and the multimedia resources in the to-be-read information to obtain a reading file, so that a playing side accesses the server side to obtain the reading file, and corresponding multimedia resources are displayed in a process of playing the reading file; wherein the sequentially organizing the audio segments and the multimedia resources to obtain the reading file comprises: in a case where the multimedia resources are pictures, generating description text of the pictures through a picture-text network, obtaining an audio segment corresponding to the description text, and sequentially organizing the audio segments corresponding to the text segments and the audio segments corresponding to the multimedia resources to obtain the reading file; in a case where the multimedia resources are videos, extracting an audio file in the videos as an audio segment, and sequentially organizing the audio segments corresponding to the text segments and the audio segments corresponding to the multimedia resources to obtain the reading file.
2. The method of claim 1, wherein, The segmenting the text in the to-be-read information to obtain the plurality of text segments comprises: traversing the to-be-read information, continuously collecting text until a multimedia resource is detected or the traversal ends in a case where the text is detected, and taking the currently collected text as a text segment; in a case where the multimedia resource is detected, associating the multimedia resource with a recently collected text segment; The sequentially organizing the audio segments and the multimedia resources to obtain the reading file comprises: adding the associated multimedia resource at a playing end position of each audio segment, and integrating to obtain the reading file.
3. The method of claim 1, wherein, The method further comprises: for each multimedia resource, determining at least one text segment adjacent to the multimedia resource; obtaining a picture-text relevance degree of each sentence in the multimedia resource and the adjacent text segment, and determining a relevant sentence with the highest picture-text relevance degree therefrom; The sequentially organizing the audio segments and the multimedia resources to obtain the reading file comprises: for an audio segment corresponding to the adjacent text segment, determining an audio playing position of the relevant sentence; adding the multimedia resource at the audio playing position of the relevant sentence corresponding to the multimedia resource, and integrating to obtain the reading file.
4. The method according to claim 2 or 3, characterized in that, The sequentially organizing the audio segments and the multimedia resources to obtain the reading file comprises: setting a unique identifier for each text segment; organizing the audio segment corresponding to the text segment, the added multimedia resource and a resource type thereof into a data slice together with the unique identifier; sequentially integrating all the data slices according to a position sequence of the text segments to form a complete reading file.
5. The method of claim 1, wherein, The determining the to-be-read information according to the indication message comprises: determining a content type of the indication message; the content type comprises at least any one of a webpage link and information text; parsing the indication message according to the content type to determine the to-be-read information.
6. The method of claim 5, wherein, The determining the content type of the indication message comprises: obtaining a field value representing a message type in a transmission encapsulation format of the indication message; in a case where the field value represents a link, determining that the content type is a web link; in a case where the field value represents text, detecting whether a string in the indication message includes a link in text form by using a regular expression, and if yes, determining that the content type is a web link, and if no, determining that the content type is information text; in a case where the field value represents an image, extracting text information from the image, and detecting whether the text information belongs to a web link or information text by using a regular expression.
7. The method of claim 5, wherein, The parsing the indication message according to the content type to determine the information to be read comprises: in a case where the content type is a web link, locating to web content according to the link in the indication message, parsing a web structure of the web content, and sequentially obtaining web text and multimedia resources as the information to be read according to a relative position relationship between the web text and the multimedia resources; or, in a case where the content type is information text, obtaining a field value representing message content in a transmission encapsulation format of the indication message as the information to be read.
8. The method of claim 1, wherein, The server comprises a custom server on a WeChat platform service; and the receiving the indication message from the user terminal comprises: receiving the indication message forwarded by a WeChat server; and the WeChat server receives the indication message sent by a user through a client of the WeChat platform service.
9. A cross-device information broadcasting device, characterized in that, The device applied to the server comprises: an information obtaining module configured to receive the indication message from the user terminal and determine the information to be read according to the indication message; the indication message comprises multi-modal information for locating the information to be read; a text segmentation processing module configured to segment text in the information to be read according to a position of a multimedia resource in the information to be read to obtain a plurality of text segments; a reading file constructing module configured to convert the text segments into corresponding audio segments, and sequentially organize the audio segments and the multimedia resource to obtain a reading file according to a position order of the text segments and the multimedia resource in the information to be read, so that the playing terminal accesses the server to obtain the reading file and displays the corresponding multimedia resource in the process of playing the reading file; wherein the sequentially organizing the audio segments and the multimedia resource to obtain the reading file comprises: in a case where the multimedia resource is an image, generating description text of the image through a graphic-text network, obtaining an audio segment corresponding to the description text, and sequentially organizing the audio segments corresponding to the text segments and the audio segments corresponding to the multimedia resource to obtain the reading file; in a case where the multimedia resource is a video, extracting an audio file in the video as an audio segment, and sequentially organizing the audio segments corresponding to the text segments and the audio segments corresponding to the multimedia resource to obtain the reading file.
10. A cross-device information casting system, comprising: The system comprises a user terminal, a server and a playing terminal; wherein The user terminal sends an instruction message to the service terminal, so that the service terminal determines to-be-read information according to the instruction message; the instruction message comprises multi-modal information used for locating the to-be-read information; The service terminal segments text in the to-be-read information according to positions of multimedia resources in the to-be-read information, to obtain a plurality of text segments, converts the text segments into corresponding audio segments, and sequentially organizes the audio segments and the multimedia resources according to a position sequence of the text segments and the multimedia resources in the to-be-read information, to obtain a read file; The playing terminal accesses the service terminal to obtain the read file, and displays corresponding multimedia resources in a process of playing the read file; The sequentially organizing the audio segments and the multimedia resources to obtain the read file comprises: In a case where the multimedia resource is a picture, description text of the picture is generated by using a picture-text network, an audio segment corresponding to the description text is obtained, and the audio segment corresponding to the description text and an audio segment corresponding to the multimedia resource are sequentially organized to obtain the read file; In a case where the multimedia resource is a video, an audio file in the video is extracted as an audio segment, and the audio segment corresponding to the text segment and the audio segment corresponding to the multimedia resource are sequentially organized to obtain the read file.
11. An electronic device, comprising: Comprise: a memory and a processor; the memory is configured to store a computer program; the processor is configured to call the computer program to implement the method according to any one of claims 1-8.
12. A readable storage medium, having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the method according to any one of claims 1-8.
Citation Information
Patent Citations
Data processing method and device, storage medium and terminal
CN111263058A
Sound and text synchronization broadcasting method and broadcasting system
CN115580742A