Cross-equipment screen projection control method and device, medium and vehicle
By introducing an artificial intelligence model to recognize the video playback status on the vehicle side and using the Android SDK to simulate click operations, the problem of video applications that do not integrate the manufacturer's voice SDK being unable to respond to voice commands on the vehicle side has been solved, achieving wider applicability and stability of screen projection control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-03-10
AI Technical Summary
In existing technologies, video applications that do not integrate the manufacturer's voice SDK cannot respond to voice commands from the vehicle, resulting in insufficient applicability and stability of vehicle-side screen projection control.
By introducing an artificial intelligence model to identify the playback status of video on the vehicle, and using the simulated click method provided by the Android SDK, the system can control the video on the vehicle and switch the playback status.
Without relying on the mobile terminal voice SDK, effective control of in-vehicle screen projection was achieved, improving the applicability and stability of the screen projection control method.
Smart Images

Figure CN121644860A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent interaction, and in particular to a cross-device screen projection control method and device, medium and vehicle. BACKGROUND
[0002] With the continuous improvement of people's demand for multi-source and intelligent functions of vehicles, vehicle and mobile terminal interconnection technology has emerged and matured. As a typical application of vehicle and mobile terminal interconnection technology, vehicle projection technology has been widely deployed in various vehicle scenes. After projecting the video of the mobile terminal to the display screen of the vehicle through the vehicle projection technology, in order to improve the user's interactive experience and the convenience of operation, the vehicle-mounted system of the vehicle needs to be able to control the current video playing state based on the triggering of the user's instruction.
[0003] In related technologies, the response to the voice instruction of the vehicle mainly depends on the voice software development kit (Software Development Kit, SDK) of the mobile terminal manufacturer. Specifically, after receiving the voice instruction of the user, the vehicle transmits the voice instruction to the mobile terminal through the interconnection link, and switches the playing state of the video based on the voice SDK embedded in the mobile terminal, to realize the response to the voice instruction of the user. However, there are various video applications on the market at present, but only a small number of video applications access the voice SDK of the mobile terminal manufacturer, resulting in that the video applications that do not integrate the voice SDK of the manufacturer cannot realize the response to the voice instruction from the vehicle through the above-mentioned method.
[0004] Therefore, how to realize effective control of the vehicle projection without relying on the voice SDK of the mobile terminal, and respond to the voice instruction of the user from the vehicle, has become a technical problem to be solved at present. SUMMARY
[0005] Therefore, the embodiments of the present application provide a cross-device screen projection control method, device, medium and vehicle, which can identify the playing state of the vehicle video through an artificial intelligence model, and further use the common simulated click method provided by the Android SDK to perform click operation on the video of the vehicle, to realize the response to the voice instruction of the user.
[0006] In a first aspect, the embodiments of the present application provide a cross-device screen projection control method, which comprises: receiving a voice instruction sent by a user, and performing semantic recognition on the voice instruction; when it is recognized that the voice instruction is an instruction for controlling the screen projection of the vehicle, recognizing a video region in an interface image of a display screen of the vehicle, and determining the current playing state of the video in the video region; and based on the video region and the current playing state of the video, performing a simulated click operation in the video region to switch the current state of the video.
[0007] As a possible implementation form of the first aspect, when the semantic instruction is to switch the video to the first state, the video region in the interface image of the display screen of the vehicle end is identified, and the current playing state of the video in the video region is determined, including: continuously capturing multiple frames of images of the display screen of the vehicle end within a preset time threshold; extracting image features of each frame of image in the multiple frames of images by the first video state recognition model, and calculating the difference degree between the image features of adjacent frames of images in the multiple frames of images; and determining the video region in the interface image of the display screen of the vehicle end according to the difference degree, and determining the current playing state of the video in the video region.
[0008] As a possible implementation form of the first aspect, when the semantic instruction is to switch the video to the second state, the video region in the interface image of the display screen of the vehicle end is identified, and the current playing state of the video in the video region is determined, including: capturing a frame of image of the current display screen of the vehicle end; obtaining an image of the display screen recorded at the time closest to the current time when the video is switched to the first state, and taking the image as a target image; extracting image features of the current image and the target image respectively by the second video state recognition model, and calculating the matching degree between the image features of the current image and the target image; determining the current playing state of the video based on the matching degree.
[0009] As a possible implementation form of the first aspect, before the simulated clicking operation is performed in the video region, the method further includes: determining whether there is a first region in the video region by the semantic recognition model, wherein the first region is an additional media region in the video region which does not belong to the main video content; if there is a first region in the video region, determining the position information of the first region, and the simulated clicking operation is not performed in the first region.
[0010] As a possible implementation form of the first aspect, the simulated clicking operation in the video region includes: obtaining coordinate information of at least one corner of four corners of a projection window of the display screen of the vehicle end and size information of the projection window; based on the coordinate information and the size information, calculating coordinate information of four corners of the video region by a position determination model, and determining coordinate information of a central position of the video region by a video state recognition model according to the coordinate information of the four corners of the video region, wherein the position determination model is used to represent a mapping relationship between the position of the video region and the position of the projection window; and performing the simulated clicking operation at the central position of the video region according to the coordinate information of the central position, to switch the current state of the video.
[0011] As a possible implementation manner of the first aspect, after performing the simulated click operation at the center position of the video area, the method further includes: determining whether the current state of the video is successfully switched; if not, identifying a button area in the video area through a video state identification model; after identifying the button area, determining a type of the button area through a button identification model, and determining a target button area based on the type of the button area, where the target button area is a button area corresponding to the state to which the video is to be switched; and performing a simulated click operation on the target button area to switch the current state of the video.
[0012] As a possible implementation manner of the first aspect, the simulated click operation includes: calling a simulated user operation method provided by an Android software development kit to perform the simulated click operation.
[0013] In the second aspect, an embodiment of the present application provides a cross-device screen projection control device, which includes: an instruction receiving module configured to receive a voice instruction sent by a user and perform semantic recognition on the voice instruction; a state identification module configured to, when it is identified that the voice instruction is an instruction for controlling screen projection of a vehicle end, identify a video area in an interface image of a display screen of the vehicle end through a video state identification model, and determine a current playing state of a video in the video area; and a state switching module configured to, based on the video area and the current playing state of the video, perform a simulated click operation in the video area to switch the current state of the video.
[0014] In the third aspect, an embodiment of the present application provides a computer readable storage medium, which is characterized by storing a computer program, and the computer program is configured to execute the cross-device screen projection control method in the first aspect.
[0015] In the fourth aspect, an embodiment of the present application provides a vehicle, which includes: a processor; and a memory configured to store processor-executable instructions, where the processor is configured to execute the cross-device screen projection control method in the first aspect.
[0016] The embodiments of the present application provide a cross-device screen projection control method, device, medium and vehicle. After receiving a voice instruction related to control of screen projection content of a vehicle end by a user, an image of the display screen of the vehicle end is actively captured and analyzed, the video area in the interface image of the display screen of the vehicle end and the playing state of the video are identified, the playing state is switched by performing a simulated click operation in the video area based on the identified result without relying on a related voice protocol (for example, a voice software development kit) of a mobile terminal, the response to the voice instruction of the user is realized, the dependence on a specific application and a specific protocol is eliminated, and the applicability and stability of the screen projection control method are improved. BRIEF DESCRIPTION OF DRAWINGS
[0017] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the embodiments of the present disclosure to explain the disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the description of detailed exemplary embodiments with reference to the accompanying drawings.
[0018] FIG. 1 This is a schematic diagram of the structure of a cross-device screen projection control system provided in an exemplary embodiment of this application.
[0019] FIG. 2 This is a flowchart illustrating a cross-device screen projection control method provided in an exemplary embodiment of this application.
[0020] FIG. 3 This is a flowchart illustrating a method for identifying a video area and playback status on a vehicle-mounted display screen, provided by an exemplary embodiment of this application.
[0021] FIG. 4 This is a flowchart illustrating another method for identifying a video area and playback status on a vehicle-mounted display screen, provided by an exemplary embodiment of this application.
[0022] FIG. 5 This is a schematic diagram of the projection window of a vehicle display screen provided in an exemplary embodiment of this application.
[0023] FIG. 6 This is an exemplary embodiment provided in this application. FIG. 5 The diagram shows the structure of the video area in the screen mirroring window.
[0024] FIG. 7 This is a schematic diagram of the cross-device screen projection control device provided in an exemplary embodiment of this application.
[0025] FIG. 8 This is a block diagram of an electronic device for cross-device screen projection control provided in an exemplary embodiment of this application. Detailed Implementation
[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0027] SUMMARY Vehicle-to-mobile terminal interconnection technology is an interactive technology that establishes a data transmission channel between a mobile terminal and a vehicle system through wired or wireless communication, enabling deep integration in terms of functional collaboration, data interaction, and service convergence. Vehicle-to-screen projection technology, as a typical application of vehicle-to-mobile terminal interconnection technology, is used to map the content of mobile terminal applications (e.g., video applications) to the vehicle and display it on the vehicle's electronic screen, providing users with a better driving experience.
[0028] To further enhance the user's interactive experience and ease of operation, after mapping the content of mobile terminal applications to the vehicle's electronic screen, the vehicle's in-vehicle system needs to be able to control the application based on the user's control commands. For example, in response to the user's voice commands, the system can switch the playback status of video applications.
[0029] In related technologies, vehicle-mounted screen projection control technology mainly relies on accessing the manufacturer's voice SDK in the mobile terminal to respond to voice commands from the vehicle user.
[0030] A voice SDK is a collection of pre-packaged software tools, code libraries, API interfaces, and documentation provided by voice technology providers. By partnering with manufacturers to integrate the corresponding voice SDK into mobile terminals, mobile terminals can respond to voice commands from the vehicle and switch the playback status of video applications projected onto the screen. However, while there are many types of video applications on the market, only a small fraction of them integrate mobile terminal manufacturers' voice SDKs. This means that a large number of video applications without integrated manufacturer voice SDKs cannot respond to voice commands from the vehicle and switch the playback status of projected video applications using the aforementioned method.
[0031] To address the aforementioned issues, this application's embodiments creatively introduce an artificial intelligence model to intelligently identify the playback status and video area of the video projected onto the vehicle's screen. Furthermore, based on the user's voice commands, a simulated click operation is executed within the video area using a method in the Android SDK that programsmatically simulates user click operations, switching the playback status of the video application, thereby achieving a response to the user's voice commands on the vehicle's screen.
[0032] The following will refer to the appendix. FIGS. 1-8 The following describes various non-limiting embodiments of this application.
[0033] Example Voice-Command-Based Cross-Device Casting Control System FIG. 1 This is a schematic diagram of the structure of a cross-device screen projection control system provided in some embodiments of this application. For example... FIG. 1As shown, the cross-device projection control system 100 may include at least one mobile terminal 110 and a vehicle 120. The vehicle 120 may include a control system 121 and a display device 122.
[0034] In some embodiments, the control system 121 may be a vehicle-side infotainment system for managing the connection between the vehicle and the mobile terminal, and for receiving and forwarding user instructions to control the playback status of video applications projected onto the vehicle.
[0035] Display device 122 is used to simultaneously present the content of a video application projected from mobile terminal 110 to vehicle 120 in a visual manner. For example, the display device may be a vehicle's display screen.
[0036] In some embodiments, at least one mobile terminal 110 can be connected to the vehicle 120 via a wired connection or a wireless connection. For example, at least one mobile terminal 110 and the vehicle 120 can establish a communication connection to achieve interconnection through a third-party screen mirroring application (e.g., CarPlay, CarLife) installed on the mobile terminal 110, or they can establish a communication connection wirelessly via WIFI or Bluetooth.
[0037] In some embodiments, at least one mobile terminal 110 projects a video application onto the display screen of the vehicle terminal 120, where it is displayed.
[0038] In some embodiments, the display device 122 may include at least one display screen. For example, the display device 122 may include a central display screen and an entertainment screen designed for the front passenger.
[0039] In some embodiments, at least one mobile terminal 110 may include any one or more of a mobile phone terminal, a tablet computer terminal, a laptop computer terminal, a personal computer terminal, and a wearable smart device (e.g., a smartwatch, a smart bracelet).
[0040] In some embodiments, the control system 121 may include a voice recognition module and a video status recognition module, wherein the voice recognition module is used to recognize user voice commands on the vehicle side, and the video status recognition module is used to determine the playback status of the video application currently presented on the display device 122 based on the screenshot of the display device 122 interface.
[0041] In some embodiments, the vehicle terminal 120 can execute the cross-device screen projection control method described below through the control system 121.
[0042] Example Cross-Device Casting Control Method To further explain FIG. 1The application illustrates the specific process of controlling screen projection using a cross-device projection control system. It also provides an exemplary flowchart of a cross-device projection control method. FIG. 2 This cross-device projection control method is applied to the cross-device projection control system 100.
[0043] like FIG. 2 As shown, the control system 121 on one side of vehicle 120 can perform the following steps: S210: Receive voice commands sent by the user and perform semantic recognition on the voice commands.
[0044] Voice commands are instructions sent by the vehicle user to control the playback status of video applications displayed on the vehicle's screen. Specifically, voice commands can be "pause playback," "start playback," "continue playback," etc.
[0045] In some embodiments, a vehicle-mounted voice system can be used to collect the user's voice information to recognize the user's voice commands. For example, the vehicle-mounted voice system may be a sensor (e.g., a microphone array) built into the vehicle. The vehicle-mounted voice system collects the user's voice information to further enable voice interaction between humans and computer systems.
[0046] In some embodiments, after collecting user voice information through an in-vehicle voice system, a trained semantic recognition model can be used to extract semantics from the collected voice information. This allows for further automatic analysis of the voice information, identifying key voice information related to controlling video playback status, and accurately recognizing user commands. For example, it can automatically recognize messages such as "pause video playback" and "start video playback."
[0047] In some embodiments, the semantic recognition model can be an artificial intelligence model, specifically, it can be a traditional machine learning model or an end-to-end deep learning network model, without any specific limitation.
[0048] It should be noted that the method for recognizing voice commands provided in the above embodiments is merely an example, and other methods can also be used to recognize user voice commands. For example, keyword templates can be directly embedded in the system, and the voice commands can be converted into semantic information through a semantic extraction model. Then, the semantic information can be matched with the keywords in the template to recognize the voice commands that control the video playback status. This application does not specifically limit the method of voice recognition.
[0049] S220. When the voice command is recognized as a command to control the screen projection on the vehicle, the video area in the interface image of the display screen on the vehicle is identified, and the current playback status of the video in the video area is determined.
[0050] In some embodiments, a video region in the vehicle-mounted display screen interface image can be identified using a video state recognition model.
[0051] For example, the video state recognition model can be an image recognition model, such as the YOLO model, the Qwen3-VL model, etc.
[0052] Considering that the process of determining the playback status of the video displayed on the vehicle screen varies depending on the user's semantic commands, this application uses different methods to identify the playback status of the video on the vehicle for the "play" voice command and the "pause" semantic command.
[0053] In some embodiments, the video state recognition model includes a first video state recognition model and a second video state recognition model, wherein the first video state recognition model is mainly used to identify whether the video is in a playing state, and the second video state recognition model is mainly used to identify whether the video is in a paused state. For example, the video state recognition model can be an image recognition model.
[0054] As an example of recognizing a video as being in a playback state, when the voice command in S210 is recognized as a pause command, meaning the video is switched from its current state to a first state, where the first state can be a paused state, it is necessary to confirm whether the current video is in a playback state. FIG. 3 As shown, S220 may include the following sub-steps: S221. Within a preset time threshold, continuously capture multiple frames of images from the vehicle's display screen.
[0055] In some embodiments, the time threshold can be 0.5s. For example, 10 frames of images from the vehicle display screen can be captured within 0.5s.
[0056] It should be noted that the time threshold of 0.5s and the number of images captured are merely examples. In actual applications, they can be flexibly set according to actual needs. For instance, if the changes in the video playback during the current time period are small, the time threshold can be increased accordingly, and the number of images captured from the vehicle display screen can also be increased. For example, the time threshold can be set to 2s, and the number of images captured from the vehicle display screen can be increased to 20.
[0057] In some embodiments, a timestamp can be added to the image captured in the aforementioned S221 so that the captured image of the vehicle display screen contains time information, unifying the global time source, thereby improving the timing accuracy and synchronization of the image.
[0058] It should be noted that time information can also be provided to the captured images of the vehicle display screen in other ways. For example, time information can be added to the image file name or high-precision, consistent time information can be provided to the image through time synchronization protocols (e.g., Precision Time Protocol (PTP) and Network Time Protocol (NTP)). This application does not specifically limit the method of providing time information for the image.
[0059] S222. Extract the image features of each frame in the multi-frame images using the first video state recognition model, and calculate the difference between the image features of adjacent frames in the multi-frame images.
[0060] S223. Determine the video area in the interface image of the vehicle-mounted display screen based on the degree of difference, and determine the current playback status of the video in the video area.
[0061] The first video state recognition model can be an image recognition model that determines the current playback state of a video by analyzing the temporal changes and content differences between consecutive video frames.
[0062] The video area refers to the area on the vehicle's display screen where video content projected from the mobile terminal is displayed.
[0063] In some embodiments, the first video state recognition model may include two parts: a feature extraction module and a feature analysis module. The specific recognition process is as follows: the multiple frames of images captured in S221 are input into the video state recognition model. The first video state recognition model extracts features from each input frame image through the feature extraction module, and then further analyzes the extracted image features through the feature analysis module to obtain the playback state of the video corresponding to the captured image. For example, the feature analysis module can calculate the difference between features of adjacent frames to analyze the temporal changes in the video, thereby determining the video area in the vehicle display screen interface and the current playback state of the video presented in that video area.
[0064] In some embodiments, the first video state recognition model can be an artificial intelligence model, such as a deep learning network model. By training the first video state recognition model with training samples, it learns the relationship between the degree of difference in changes between adjacent frame images and the video state. This enables the video state recognition model to automatically identify the playback state of the video corresponding to a given frame based on consecutive video frame images, thereby improving the accuracy of video playback state recognition.
[0065] It should be noted that the video playback status can also be determined by calculating the difference between adjacent image frames in other ways. For example, the difference between adjacent image frames can be calculated based on feature vectors. For instance, the Euclidean distance, cosine distance, or Manhattan distance between image features can be calculated to determine the difference between adjacent image features. Alternatively, the difference between adjacent image frames can be calculated based on the image content. For example, the Structural Similarity Index (SSIM) can be calculated, combining brightness, contrast, and structure to determine the difference. This application does not specifically limit the calculation of the difference between adjacent image frames.
[0066] As an example, based on the triggering of the pause voice command, 10 frames of the vehicle's display screen interface can be captured continuously within 0.5 seconds. These 10 frames are then input into the video state recognition model, which directly outputs the current playback state of the video corresponding to these 10 frames as "playing".
[0067] It should be noted that the aforementioned S221-S223 can also be used to identify when a video is paused, and no specific limitation is made here.
[0068] In some embodiments, the multiple frames of images captured in S221 can be stored in the vehicle's storage device, such as memory, to provide a data basis for subsequent identification of the video playback status.
[0069] Considering that determining whether a video is in a "paused" state only requires comparing the current screenshot with the closest frame from the last time the video was playing, as an example of recognizing a paused video, when the voice command in S210 is recognized as a play command, i.e., switching the video to the second state, where the second state could mean the video is playing, it is necessary to confirm whether the current video is paused. FIG. 4 As shown, S220 may include the following sub-steps: S224. Capture a frame of the current vehicle-mounted display screen.
[0070] S225. Obtain the saved image of the display screen when the video closest to the current time switches to the first state, and use that image as the target image.
[0071] In some embodiments, when the video state is switched to the first state (paused state) in response to a user's voice command, an image of the vehicle-mounted display screen is captured and recorded to provide data support for subsequent determination of the video state.
[0072] The target image can be the image of the display screen that is closest to the current moment when the video was switched to the first state, that is, the image of the display screen captured when the video was switched to the pause state in response to the user's voice command.
[0073] S226. Extract the image features of the current image and the target image respectively through the second video state recognition model, and calculate the matching degree between the image features of the current image and the target image.
[0074] S227. Determine the current playback status of the video based on the matching degree.
[0075] The second video state recognition model can be an image recognition model that determines the current playback state of a video by calculating the degree of matching between the current image frame and the target image.
[0076] In some embodiments, the second video state recognition model is similar to the first video state recognition model, comprising two parts: a feature extraction module and a matching degree calculation module. The specific recognition process is as follows: After selecting the target image, the captured image of the current interface of the vehicle's display screen and the target image are input into the second video state recognition model. The second video state recognition model extracts features from both the target image and the screenshot of the current interface of the vehicle's display screen using the feature extraction module. Furthermore, the matching degree calculation module calculates the matching degree (or similarity) between the features of the target image and the features of the screenshot of the current interface of the vehicle's display screen. Based on this matching degree, the current playback state of the video is determined, primarily whether the video is currently paused. For example, when the matching degree is greater than or equal to 95%, it can be determined that the video is currently paused.
[0077] S230. Based on the video area and the current playback state of the video, perform a simulated click operation within the video area to switch the current state of the video.
[0078] Simulated clicks can be a method that uses programming to simulate user clicks on a vehicle's display screen. For example, a public click method provided by the View class in the Android application development kit can be used to simulate clicks, such as the performClick() method. This method programmatically triggers a click event listener (OnClickListener) to simulate the user's click behavior.
[0079] In some embodiments, after determining the video region using a video state recognition model, the coordinate information of the video region is further output. Specifically, the coordinate information of at least one of the four corners of the projection window of the vehicle's display screen and the window's size information can be pre-input. Based on the coordinate information and size information of the projection window, the coordinate information of the four corners of the video region in the projection window is determined using a position determination model. The position determination model is used to characterize the mapping relationship between the position of the video region and the position of the projection window. Furthermore, based on the coordinate information of the four corners of the video region, the coordinate information of the center position of the video region is determined, and a simulated click operation is performed at the center position to achieve the switching of the video state.
[0080] In some embodiments, the location determination model can be a deep learning model, which is trained to learn the relationship between the location of a video region and the projection window by inputting sample images into the location determination model.
[0081] By utilizing the common simulated click method provided in the Android software development kit (SDK), simulated user clicks are performed within the area of the content projected onto the vehicle's display screen. This enables the switching of the playback status of video applications projected onto the vehicle, eliminating the reliance on the mobile terminal's voice software development kit and thus removing limitations on the type of video application. This allows the projection control method to respond to user voice commands on the vehicle, enabling control of various types of video applications projected onto the vehicle, significantly improving the versatility and applicability of the projection control method.
[0082] As an example, such as FIG. 5 As shown, the coordinates of the upper left corner of the projection window 510 in the vehicle's display screen 500 are T(2, -2). The coordinates of point T are in a coordinate system with the upper left corner of the vehicle's display screen 500 as the origin. The projection window has a length of 32 and a width of 18, with a length-to-width ratio of 16:9. Based on the coordinates and dimensions of this projection window, the coordinates of the four corners of the video area 511 are determined using a position determination model: A(4, -4), B(20, -4), C(4, -20), and D(20, -20). Furthermore, the coordinates of the center position of the video area are calculated as (12, -12) using the coordinates of the four corners. A simulated click operation is performed at the coordinate position (12, -12) to switch the video state.
[0083] It should be noted that the above is merely an example of calculating the coordinates of the center position of the video area. You can also input the coordinates of the upper right corner, lower left corner, or lower right corner of the projection window. You can also input the coordinates of two corners of the video area at the same time, such as the upper left corner and the upper right corner. You can also input the coordinates of all four corners of the projection window at the same time. This application does not impose any specific limitations on this.
[0084] Considering that in practical applications, after performing a simulated click on the video area, further video switching function buttons may appear in the video area, such as a pause button or a play button, requiring further clicks to switch the video state. Therefore, after performing a simulated click on the video area, it is necessary to further confirm whether the current video state has been successfully switched. Specifically, the current video playback state can be confirmed using the aforementioned method for identifying the video playback state (S221-S223 or S224-S226), combined with the user's semantic command to determine whether the current video state has been successfully switched. For example, when the user's voice command is a pause command, the aforementioned video state recognition method can be used to confirm whether the current video playback state is paused. If it is paused, the video playback state is considered to have switched successfully. For methods for identifying video playback state, please refer to [link to relevant documentation]. FIG. 3 , FIG. 4 The relevant descriptions of its embodiments are not repeated here.
[0085] In some embodiments, when it is confirmed that the current video state has not been successfully switched, a video recognition model can be used to identify button areas within the video area. Specifically, after identifying a button area, the button recognition model can be used to further determine the type of the button area, and a target button area can be determined based on the type of the button area. The target button area can be a button area corresponding to the state to which the video needs to be switched. After confirming the target button area, a simulated click operation is performed again within the target button area to switch the current playback state of the video.
[0086] In some embodiments, the button recognition model can be a deep learning network model, for example, an object detection model, such as the YOLO model, R-CNN model, etc.
[0087] As an example, such as FIG. 6As shown, the button area 511 in the video area 510 may include a video state switching function button 5111 and a progress adjustment button 5112. A button recognition model can be used to identify the position of the video state switching function button 5111 in the button area and the specific type of video state that the function button switches. For example, if the video needs to be switched to a paused state, it is determined whether the video state switching function button is a pause button. If so, the video state switching function button is used as the target button area, and the coordinates of the center position of the target button area are further determined. A simulated click operation is then performed at the center position of the target button area to switch the video playback state to paused.
[0088] Considering that interference factors such as advertisements and pop-ups, which are not part of the main video content, may exist during the process of determining whether a video is in a playing state, affecting the accuracy of the video state recognition model in identifying the video playback state, in some embodiments, after identifying the video region and the video playback state, a semantic recognition model is further used to determine whether a first region exists in the video. The first region can be an additional media region within the video region that is not part of the main video content being displayed; for example, the first region can be an advertisement or a pop-up within the video region. After determining that the first region exists in the video region, the video state recognition model further determines the location information of the first region, such as the coordinates of the four corners of an advertisement, to exclude additional media regions that are not part of the main video content being displayed, and to prevent simulated click operations from being performed within these additional media regions, thereby improving the accuracy of video playback state recognition and reducing the overall error rate.
[0089] As an example, after identifying a video region using a video state recognition model, a semantic recognition model can be used to extract text from the screenshot of that region and identify keywords that might be part of the first region to determine its existence. For example, keywords related to the first region could be "advertisement," "promotion," "recommendation," "claim now," or "activate now." After detecting keywords related to the first region, an image recognition model can be used to further determine the specific location information of the first region, such as the coordinates of its four corners.
[0090] It should be noted that the method of identifying the first region described in the foregoing embodiments is merely an example of identifying the first region (advertising region). The first region in the video region can also be identified through all other possible methods, and this application does not make any specific limitations on this.
[0091] In summary, the cross-device screen projection control method provided in this application actively captures and analyzes the image of the vehicle-side display screen after receiving a user's voice command related to the projection content on the control vehicle. It uses an artificial intelligence model to identify the video area and playback status of the video in the interface image of the vehicle-side display screen. Furthermore, by leveraging the common simulated click method provided in the Android software development kit, it switches the playback status by performing simulated click operations within the video area without relying on the relevant voice protocols of the mobile terminal (e.g., the voice software development kit), thus responding to the user's voice commands. This eliminates dependence on specific applications and protocols, significantly improving the universality, applicability, and stability of the screen projection control method.
[0092] Example Cross-Device Casting Control Apparatus The above text combined FIGS. 1-6 The method embodiments of this application have been described in detail above. The apparatus embodiments of this application are described in detail below. It should be understood that the descriptions of the method embodiments correspond to the descriptions of the apparatus embodiments; therefore, any parts not described in detail can be referred to the foregoing method embodiments. FIG. 7 This is a schematic diagram of a system module of a cross-device screen projection control device shown in some embodiments of this application.
[0093] like FIG. 7 As shown, the cross-device screen projection control device 700 may include an instruction receiving module 710, a status recognition module 720, and a status switching module 730. Among them, The instruction receiving module 710 can be configured to receive voice instructions sent by the user and perform semantic recognition on the voice instructions.
[0094] The status recognition module 720 can be configured to identify the video area in the interface image of the vehicle display screen and determine the current playback status of the video in the video area when the voice command is recognized as a command to control the screen projection of the vehicle.
[0095] The state switching module 730 can be configured to perform a simulated click operation within the video area based on the video area and the current playback state of the video, in order to switch the current state of the video.
[0096] In some embodiments, a vehicle-mounted voice system can be used to collect the user's voice information to recognize the user's voice commands. For example, the vehicle-mounted voice system may be a sensor (e.g., a microphone array) built into the vehicle. The vehicle-mounted voice system collects the user's voice information to further enable voice interaction between humans and computer systems.
[0097] In some embodiments, after collecting user voice information through an in-vehicle voice system, a trained semantic recognition model can be used to extract semantics from the collected voice information. This allows for further automatic analysis of the voice information, identifying key voice information related to controlling video playback status, and accurately recognizing user commands. For example, it can automatically recognize messages such as "pause video playback" and "start video playback."
[0098] In some embodiments, the semantic recognition model can be an artificial intelligence model, specifically, it can be a traditional machine learning model or an end-to-end deep learning network model, without any specific limitation.
[0099] In some embodiments, the video state recognition model includes a first video state recognition model and a second video state recognition model, wherein the first video state recognition model is mainly used to identify whether the video is in a playing state, and the second video state recognition model is mainly used to identify whether the video is in a paused state. For example, the video state recognition model can be an image recognition model.
[0100] In some embodiments, the status recognition module 720 may also be configured to: when the user's voice command is recognized as a pause command, continuously capture multiple frames of images from the vehicle's display screen within a preset time threshold.
[0101] In some embodiments, the time threshold can be 0.5s. For example, 10 frames of images from the vehicle display screen can be captured within 0.5s.
[0102] It should be noted that the time threshold of 0.5s and the number of images captured are merely examples. In actual applications, they can be flexibly set according to actual needs. For instance, if the changes in the video playback during the current time period are small, the time threshold can be increased accordingly, and the number of images captured from the vehicle display screen can also be increased. For example, the time threshold can be set to 2s, and the number of images captured from the vehicle display screen can be increased to 20.
[0103] In some embodiments, a timestamp can be added to the image captured in the aforementioned S221 so that the captured image of the vehicle display screen contains time information, unifying the global time source, thereby improving the timing accuracy and synchronization of the image.
[0104] In some embodiments, the state recognition module 720 may also be configured to extract image features of each frame in a multi-frame image using a first video state recognition model, and calculate the difference between image features of adjacent frames in the multi-frame image.
[0105] In some embodiments, the state recognition module 720 may also be configured to determine the video region in the interface image of the vehicle-mounted display screen and the current playback state of the video in the video region based on the degree of difference.
[0106] The first video state recognition model can be an image recognition model that determines the current playback state of a video by analyzing the temporal changes and content differences between consecutive video frames.
[0107] The video area refers to the area on the vehicle's display screen where video content projected from the mobile terminal is displayed.
[0108] In some embodiments, the first video state recognition model may include two parts: a feature extraction module and a feature analysis module. The specific recognition process is as follows: the multiple frames of images captured in S221 are input into the video state recognition model. The first video state recognition model extracts features from each input frame image through the feature extraction module, and then further analyzes the extracted image features through the feature analysis module to obtain the playback state of the video corresponding to the captured image. For example, the feature analysis module can calculate the difference between features of adjacent frames to analyze the temporal changes in the video, thereby determining the video area in the vehicle display screen interface and the current playback state of the video presented in that video area.
[0109] In some embodiments, the first video state recognition model can be an artificial intelligence model, such as a deep learning network model. By training the first video state recognition model with training samples, it learns the relationship between the degree of difference in changes between adjacent frame images and the video state. This enables the video state recognition model to automatically identify the playback state of the video corresponding to a given frame based on consecutive video frame images, thereby improving the accuracy of video playback state recognition.
[0110] In some embodiments, the state recognition module 720 may also be configured to: when the user's voice command is recognized as a playback command, capture a frame of the current vehicle display screen, and take the frame of the multiple frames captured when the video was last in the second state that is closest to the current moment as the target image.
[0111] In some embodiments, the state recognition module 720 may also be configured to: extract image features of the current image and the target image respectively through the second video state recognition model, and calculate the matching degree between the image features of the current image and the target image.
[0112] The target image can be selected from screenshots stored in history used to determine the playback status process, and is used to determine the current playback status of the video.
[0113] For example, the target image can be the frame image closest to the current moment among multiple frames captured from the last time the video was playing. In order to ensure the accuracy of video state recognition, the two frames closest to the current moment can also be selected from multiple frames captured from the last time the video was playing as the target image. This application does not specifically limit the number of target images.
[0114] In some embodiments, the state recognition module 720 may also be configured to determine the current playback state of the video based on the matching degree.
[0115] In some embodiments, the second video state recognition model is similar to the first video state recognition model, comprising two parts: a feature extraction module and a matching degree calculation module. The specific recognition process is as follows: After selecting the target image, the captured image of the current interface of the vehicle's display screen and the target image are input into the second video state recognition model. The second video state recognition model extracts features from both the target image and the screenshot of the current interface of the vehicle's display screen using the feature extraction module. Furthermore, the matching degree calculation module calculates the matching degree (or similarity) between the features of the target image and the features of the screenshot of the current interface of the vehicle's display screen. Based on this matching degree, the current playback state of the video is determined, primarily whether the video is currently paused. For example, when the matching degree is greater than or equal to 95%, it can be determined that the video is currently paused.
[0116] In some embodiments, the state recognition module 720 may also be configured to perform a simulated click operation within the video area based on the video area and the current playback state of the video, so as to switch the current state of the video.
[0117] In some embodiments, after determining the video region using a video state recognition model, the coordinate information of the video region is further output. Specifically, the coordinate information of at least one of the four corners of the projection window of the vehicle's display screen and the window's size information can be pre-input. Based on the coordinate information and size information of the projection window, the coordinate information of the four corners of the video region in the projection window is determined using a video state recognition model. Furthermore, based on the coordinate information of the four corners of the video region, the coordinate information of the center position of the video region is determined, and a simulated click operation is performed at the center position to achieve the switching of video states.
[0118] By utilizing the common simulated click method provided in the Android software development kit (SDK), simulated user clicks are performed within the area of the content projected onto the vehicle's display screen. This enables the switching of the playback status of video applications projected onto the vehicle, eliminating the reliance on the mobile terminal's voice software development kit and thus removing limitations on the type of video application. This allows the projection control method to respond to user voice commands on the vehicle, enabling control of various types of video applications projected onto the vehicle, significantly improving the versatility and applicability of the projection control method.
[0119] In some embodiments, the state recognition module 720 is further configured to: confirm the current playback state of the video using the aforementioned method for recognizing the video playback state, and determine whether the current video state has been successfully switched by combining the user's semantic command. For example, when the user's voice command is a pause voice command, the aforementioned video state recognition method can be used to confirm whether the current video playback state is paused. If it is paused, the video playback state is considered to have switched successfully. For more information on the method for recognizing video playback state, please refer to [link to relevant documentation]. FIG. 3 , FIG. 4 The relevant descriptions of its embodiments are not repeated here.
[0120] In some embodiments, the video state recognition module 720 can also be configured to: when it is confirmed that the current video state has not been successfully switched, identify the button area in the video area using a video recognition model. Specifically, after identifying the button area, the type of the button area can be further determined using a button recognition model, and a target button area can be determined based on the type of the button area. The target button area can be a button area corresponding to the state to be switched in the video. After confirming the target button area, a simulated click operation is performed again within the target button area to switch the current playback state of the video.
[0121] In some embodiments, the video state recognition module 720 can also be configured to: after recognizing the video region and the video playback state, further determine whether a first region exists in the video using a semantic recognition model. The first region can be an additional media region within the video region that does not belong to the main video content being displayed; for example, the first region can be an advertisement or pop-up window within the video region. After determining that a first region exists in the video region, further determine the location information of the first region using the video state recognition model, such as the coordinates of the four corners of an advertisement, to exclude additional media regions that do not belong to the main video content being displayed, and to avoid performing simulated click operations within these additional media regions, thereby improving the accuracy of video playback state recognition and reducing the overall error rate.
[0122] In some embodiments, the video state recognition module 720 can also be configured to: after recognizing a video region through the video state recognition model, extract text from the screenshot of the video region through a semantic recognition model, and identify keywords that may be related to the first region to determine whether the first region exists. For example, keywords for the first region could be "advertisement," "promotion," "recommendation," "claim now," "activate now," etc. After detecting keywords related to the first region, the specific location information of the first region can be further determined through an image recognition model, such as the coordinates of the four corners of the first region.
[0123] It should be understood that specific limitations regarding the device can be found in the limitations regarding cross-device screen projection control methods described above, and will not be repeated here. Each module in the aforementioned device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0124] Example Electronic Device and Computer-Readable Storage Medium This application also provides an electronic device, such as FIG. 8 FIG. 1 FIG. 2 FIG. 3 FIG. 4 FIG. 5 FIG. 3 FIG. 4 FIG. 6 Example Cross-Device Casting Control Apparatus FIGS. 1-6 FIG. 7 FIG. 7 FIG. 3 FIG. 4 Example Electronic Device and Computer-Readable Storage Medium FIG. 8 FIG. 1 FIG. 2 As shown. The electronic device 800 provided in this application includes a memory 810, a processor 820, and an input / output interface 830. The memory 810, processor 820, and input / output interface 830 are connected via internal connection paths. The memory 810 stores instructions, and the processor 820 executes the instructions stored in the memory 810 to control the input / output interface 830 to receive input data and information, and output operation results and other data.
[0125] It should be understood that in the embodiments of this application, the processor 820 may be a general-purpose central processing unit (CPU), GPU, FPGA, microprocessor, application specific integrated circuit (ASIC), or one or more integrated circuits to execute related programs in order to implement the technical solutions provided in the embodiments of this application.
[0126] The memory 810 may include read-only memory and random access memory, and provides instructions and data to the processor 820. A portion of the processor 820 may also include non-volatile random access memory. For example, the processor 820 may also store device type information.
[0127] In implementation, each step of the above method can be completed by the integrated logic circuits in the hardware of the processor 820 or by instructions in software form. The cross-device screen projection control method disclosed in the embodiments of this application can be directly manifested as being executed by the hardware processor, or being executed by a combination of hardware and software modules in the processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory 810, and the processor 820 reads the information in memory 810 and completes the steps of the above method in combination with its hardware. To avoid repetition, it will not be described in detail here. This application also provides a computer program product, including a computer program / instructions. When the computer program / instruction processor in the computer program product provided in this application is executed, the cross-device screen projection control method provided in this application can be implemented.
[0128] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this application, and will not be described in detail here.
[0129] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0130] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0131] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0132] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0133] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0134] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program verification codes, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0135] It should be noted that in the description of this application, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0136] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications or equivalent substitutions made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A cross-device screen projection control method, applied to a vehicle, comprising: receiving a voice instruction sent by a user, and performing semantic recognition on the voice instruction; when it is recognized that the voice instruction is an instruction for controlling screen projection at a vehicle end, recognizing a video region in an interface image of a display screen of the vehicle end, and determining a current playing state of a video in the video region; and based on the video region and the current playing state of the video, performing a simulated click operation in the video region to switch the current state of the video. When the semantic instruction is to switch the video to a first state, recognizing the video region in the interface image of the display screen of the vehicle end, and determining the current playing state of the video in the video region comprises:
2. The cross-device screen projection control method according to claim 1, characterized in that, within a preset time threshold, continuously capturing multiple frames of images of the display screen of the vehicle end; extracting image features of each frame of image in the multiple frames of images by the first video state recognition model, and calculating a difference degree between image features of adjacent frames of images in the multiple frames of images; and determining the video region in the interface image of the display screen of the vehicle end according to the difference degree, and determining the current playing state of the video in the video region. When the semantic instruction is to switch the video to a second state, the recognizing the video region in the interface image of the display screen of the vehicle end, and determining the current playing state of the video in the video region comprises:
3. The cross-device screen projection control method according to claim 1, characterized in that, capturing a frame of image of the current display screen of the vehicle end; obtaining a saved image of the display screen when the video is switched to the first state closest to the current time, and taking the image as a target image; extracting image features of the current image and the target image respectively by the second video state recognition model, and calculating a matching degree between the image features of the current image and the target image; based on the matching degree, determining the current playing state of the video. Before performing the simulated click operation in the video region, further comprising:
4. The cross-device screen projection control method according to claim 2, characterized in that, determining whether a first region exists in the video region by a semantic recognition model, wherein the first region is an additional media region in the video region that does not belong to main video content; if the first region exists in the video region, determining position information of the first region, and not performing the simulated click operation in the first region. The performing the simulated click operation in the video region comprises:
5. The cross-device screen projection control method according to claim 2, characterized in that, obtaining coordinate information of at least one corner of four corners of a screen projection window of the display screen of the vehicle end and size information of the screen projection window; based on the coordinate information and the size information, calculating coordinate information of four corners of the video region by a position determination model, and determining coordinate information of a central position of the video region according to the coordinate information of the four corners of the video region, wherein the position determination model is used to represent a mapping relationship between a position of the video region and a position of the screen projection window; according to the coordinate information of the central position, performing the simulated click operation at the central position of the video region to switch the current state of the video. 6. The cross-device screen projection control method according to claim 5, characterized in that, The method further comprises, after performing the simulated click operation at the center position of the video region: determining whether the current state of the video is successfully switched; if not, identifying a button region in the video region by using the video state recognition model; after identifying the button region, determining the type of the button region by using a button recognition model, and determining a target button region based on the type of the button region, wherein the target button region is a button region corresponding to the state to which the video is to be switched; performing the simulated click operation on the target button region to switch the current state of the video.
7. The cross-device screen projection control method according to any one of claims 1 to 6, characterized in that, The simulated click operation comprises: calling an Android software development kit (SDK) provided simulated user operation method to perform the simulated click operation.
8. A cross-device screen projection control device, the device comprising: an instruction receiving module configured to receive a voice instruction sent by a user and perform semantic recognition on the voice instruction; a state recognition module configured to, when it is identified that the voice instruction is an instruction for controlling screen projection of the vehicle end, identify a video region in an interface image of a display screen of the vehicle end, and determine a current playing state of a video in the video region; and a state switching module configured to, based on the video region and the current playing state of the video, perform a simulated click operation in the video region to switch the current state of the video. The storage medium stores a computer program for executing the cross-device screen projection control method of any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, a processor; 10. A vehicle characterized by comprising: a memory for storing instructions executable by the processor, wherein the processor is configured to execute the cross-device screen projection control method of any one of claims 1 to 7.