Multiplayer gaming system and method

By executing one game instance locally and a second instance remotely, with synchronized audio and video processing, the method addresses the processing limitations of devices to enable local multiplayer gaming experiences for multiple users.

GB2644126APending Publication Date: 2026-03-18SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2026-03-18

AI Technical Summary

Technical Problem

The increasing complexity and visual quality of games require more powerful processing hardware, limiting the ability to provide local multiplayer experiences due to technical constraints, especially in devices with standardised hardware capabilities.

Method used

A method and system that enables local multiplayer gaming by executing one game instance locally and a second instance remotely, with data transmission and synchronization to create a unified gaming experience across multiple devices, including audio and video processing to manage latency and processing loads.

Benefits of technology

Enables local multiplayer experiences on devices with limited processing power by utilizing remote execution and synchronized audio and video processing, allowing more players to participate simultaneously than the device could otherwise support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A system for providing synchronised multi-user application experience, such as local multiplayer gameplay, at a local processing device. First and second application processing units execute respectiv
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE INVENTION Field of the invention This disclosure relates to a multiplayer gaming system and method. Description of the Prior Art The "background" description provided herein is for the purpose of generally presenting the context of the disclosure. Work of the presently named inventors, to the extent it is described in this background section, as well as aspects of the description which may not otherwise qualify as prior art at the time of filing, are neither expressly nor impliedly admitted as prior art against the present invention. With the increasing level of complexity and visual quality that are able to be provided in games, the demands upon processing hardware have increased significantly over time. For instance, more powerful processing hardware may be required to execute games, or limits may be placed upon the functionality of such games in order to enable them to be executed by a typical device. One example of this is multiplayer experiences that can only be experienced in an online or networked setting - rather than in so-called 'couch co-op' or split-screen modes within a game. While there is a desire to offer a more local experience, technical limitations of executing devices can impose a restriction. This technical limitation may be imposed by the hardware of a games console (which is largely standardised by the manufacturer), or by any other computing arrangement; for instance, content may be developed on the basis of known console capabilities or information about average or expected computing power available to eventual users. While in some cases a device may have sufficient computing power to provide a local multiplayer experience for a game (particularly in the case of a modern computing device executing an older game), many games may not be designed with this in mind. It is in the context of the above discussion that the present disclosure arises. SUMMARY OF THE INVENTION This disclosure is defined by claim 1. Further respective aspects and features of the disclosure are defined in the appended claims. It is to be understood that both the foregoing general description of the invention and the following detailed description are exemplary, but are not restrictive, of the invention. BRIEF DESCRIPTION OF THE DRAWINGS A more complete appreciation of the disclosure and many of the attendant advantages thereof will be readily obtained as the same becomes better understood by reference to the following detailed description when considered in connection with the accompanying drawings, wherein: Figure 1 schematically illustrates an entertainment system; Figure 2 schematically illustrates a networked system; Figure 3 schematically illustrates a method; Figure 4 schematically illustrates a system for executing a video game; Figure 5 schematically illustrates a first processing device; Figure 6 schematically illustrates a method for executing a video game by a first processing device; Figure 7 schematically illustrates a method for generating and outputting combined audio; Figure 8 schematically system for providing a synchronised multi-user application experience at a local processing device; and Figure 9 schematically method for providing a synchronised multi-user application experience at a local processing device. DESCRIPTION OF THE EMBODIMENTS Referring now to the drawings, wherein like reference numerals designate identical or corresponding parts throughout the several views, embodiments of the present disclosure are described. Referring to Figure 1, an example of an entertainment system 10 is a computer or console. The entertainment system 10 comprises a central processor or CPU 20. The entertainment system also comprises a graphical processing unit or GPU 30, and RAM 40. Two or more of the CPU, GPU, and RAM may be integrated as a system on a chip (SoC). Further storage may be provided by a disk 50, either as an external or internal hard drive, or as an external solid state drive, or an internal solid state drive. The entertainment device may transmit or receive data via one or more data ports 60, such as a USB port, Ethernet® port, Wi-Fi® port, Bluetooth® port or similar, as appropriate. It may also optionally receive data via an optical drive 70. Audio / visual outputs from the entertainment device are typically provided through one or more A / V ports 90 or one or more of the data ports 60. Where components are not integrated, they may be connected as appropriate either by a dedicated data link or via a bus 100. An example of a device for displaying images output by the entertainment system is a head mounted display 'HMD' 120, worn by a user 1. Interaction with the system is typically provided using one or more handheld controllers 130, and / or one or more VR controllers (130A-L,R) in the case of the HMD. Figure 2 schematically illustrates a networked system in accordance with implementations of the present disclosure. In this Figure, a local device 200 is shown in communication, via a network represented by the line, with a remote device 210. The remote device 210 may be a processing device, such as a games console or personal computer, which is typically (but not required to be) located outside of the environment of the local device 200. This may include any location such as a different room or building; in some cases, the remote device 210 may be a processing device belonging to one of the users of the local device 200 and may be located accordingly at that user's house. Rather than being a device such as a games console, the remote device 210 may be a server which provides processing functionality so as to enable execution of a game instance at that server. The remote device 210 can therefore be any suitable device for executing a game instance and communicating with the local device 200 via a network connection. The local device 200 may comprise any computing arrangement that is communicable with a display (which may be integrated with or external to a processing device) for outputting video content to users. The local device 200 is also configured to receive inputs from control devices operated by those users, and is configured to receive data via a network connection such as a local area network or the internet. An example of this is the entertainment system 10 of Figure 1, which is able to output video to a display via the A / V port 90, can be utilised with controllers such as the handheld controller 130, and can receive data via the data port 60. Examples of suitable devices include games consoles, personal computers, televisions, mobile phones, and portable gaming devices. Implementations according to the present disclosure provide the ability for two or more users to interact with separate game instances via the same local device; this enables a local multiplayer 3 experience to be provided for content for which this would otherwise not be an option for that local device - typically due to technical constraints such as limited available processing power. At least one of these instances is executed remotely to that local device, with a network connection being used to transmit images, audio, and / or data to the local device. The present disclosure refers to games as an example of an implementation to aid the clarity of the reader's understanding; however it should be appreciated that the techniques described in this disclosure can be equally applied to any other suitable application in which multiple users may wish to participate simultaneously. This may include applications such as media applications, for example, such as applications which enable users to access free viewpoint video content - each of a plurality of users may wish to view the content from their own viewpoint, thereby causing a user desire for a multi-user arrangement. Figure 3 schematically illustrates a general method in accordance with this function. While the steps are shown in a particular order, this should not be regarded as limiting; it will be appreciated that these may be performed in any suitable order, with some steps being performed substantially simultaneously. For instance, the second instance may be initiated before the first instance, or the second player may be added whilst the execution of the second instance is initiated. A step 300 comprises executing the first instance of a game at a local processing device (that is, the device with which the users are directly interacting - such as a games console in the room with them). This instance is executed in a single player mode, in that only a single player is able to provide inputs to control the execution of that instance. A step 310 comprises adding a second user to the game; this can be in response to inputs from the user of the first instance of the game, a request from the second user, and / or a combination of the two (such as an invitation-based implementation). At this stage, this may comprise associating a user profile of the second user with the game or inserting their user avatar into a game - adding the second user is not taken to mean that the second user is able to interact with the first instance, and in the case that the second instance is not currently being executed no functionality may be available to the second user initially. Adding a second user to the game means that the first game instance and the second game instance will each provide interactivity with a shared game environment - for instance, meaning that the first user's avatar and the second user's avatar are present in the same game environment. A step 320 comprises executing a second instance of the game. This is performed by a remote device (that is, not the local processing device); while this device is typically expected to be remote in the sense that it is not present in the environment of the local device, this is not a requirement and remote may be taken to mean 'separate' in that the second instance is implemented by a device which is independent of the local device. As discussed above with reference to the remote device 210 of Figure 2, the second instance may be implemented by a second games console or a cloud gaming server, for example. A step 330 comprises transmitting an output from the second instance; this output may comprise any suitable data or content as appropriate for a given implementation. For instance, data regarding an avatar's location or interactions may be output, or the results of physics simulations associated with the second user's actions within the second instance may be output to the first instance to enable elements of the second instance to be incorporated into the first instance. Alternatively, or in addition, the output from the second instance comprises video and / or audio of the second instance - such as the rendered video showing the second user's interactions with the second instance. In the case that no video is output by the second instance, the execution of the second instance may be modified so as to not render any images - this can reduce a processing burden upon the remote device, enabling implementation by a device with reduced processing power and / or improving the energy efficiency of such an arrangement. A step 340 comprises interacting with each instance, with the interactions being controlled by users operating respective control devices which each provide inputs to the local device. In the case of the second user's inputs, these are transmitted to the remote device to allow the second instance to be controlled. The local device is configured to display the results of the interactions with the two separate instances of the games; this can be achieved in any suitable manner. For example, in some cases it may be considered appropriate to present the video output of each instance in a split-screen mode such that each instance is shown in a spatially distinct manner. This may be achieved by executing the first instance of the video game locally to generate a video output, while decoding video received from the remote device which comprises the output of the second instance of the video game. Alternatively, the first instance may be updated based upon the output of the second instance so as to represent both instances. This can comprise a shared screen for both players, so that both appear to be within the same game instance - with both player characters appearing within the same camera view, for instance. For example, an object in the first instance may move in dependence upon physics simulations performed by the second instance, with the results of those simulations (or movement information for an object, for example) being output by the second instance for use by the first instance. Implementations in accordance with the method described above can therefore provide a local multiplayer gaming experience while utilising multiple gaming instances executed by different devices. While the above has been described in the context of two players being provided with a gaming experience, it is considered that this could be extended. For example, in the case that each game 5 instance supports four players, two instances could be used to provide an (up to) eight player gaming experience in the same manner. Similarly, it is considered that a greater number of instances of a game could be utilised in combination so as to provide a gaming experience for a greater number of players. Of course, a combination of the two could also be utilised in which three or more game instances each able to support two or more players is used. In any case, the result is achieved in which more players than would otherwise be able to play a game locally are able to play via a single device despite technical limitations. Figure 4 schematically illustrates a system for executing a video game in accordance with the discussion provided above. The system comprises two or more control devices 400, a first processing device 410, a second processing device 420, and a display device 430. While shown here as a part of the system, the second processing device 420 is typically located apart from the other elements and as such can be considered to be distinct from the system as an external unit which provides inputs to the system of the elements 400, 410, and 430. The two or more control devices 400 may include any suitable input devices which enable users to provide inputs to control processing; these may be the control devices 130 or 130A of Figure 1, for example. In some cases the inputs from users may not be button presses or movements of a control device; for instance, gestures or audio inputs. In such a case, the control devices 400 may be any suitable hardware which enables the capture of these inputs such as cameras and / or microphones. The two or more control devices are associated with respective users, such that at least a first control device is associated with a first user and a second control device is associated with a second user. Here, associated may be taken to mean that they are operated by that user; however in some cases further functionality may be considered such as a control device 400 being logged in or registered with a particular user's profile or the like. The first processing device 410 is configured to receive inputs from each of the two or more control devices and to execute a first instance of a video game; the functionality of this device 410 is discussed in more detail below with reference to Figure 5. In a typical implementation, the first processing device 410 may be embodied by a games console or other local processing device; however in some cases it may be preferable that a cloud gaming service is utilised to provide this functionality. In such a case, a local device may be provided which is configured to communicate with the server so as to transmit inputs received from the controllers and to receive video of gameplay for display. The first processing device 410 is further configured to output images for display to the first and second users, the images being generated in dependence upon both the first and second instances of a video game. This may be images generated by the first instance in dependence upon data output by the 6 second instance, or may include images generated by both the first and second instances. The first processing device may also be configured to output audio for at least one of the instances of the video game; in some implementations this may include outputting audio for each instance of the video game via a different respective audio channel such that each user can be provided with audio for a corresponding instance of the video game. The second processing device 420 is configured to execute the second instance of the video game responsive to inputs from the second control device, with these inputs being provided to the second processing device 420 via the first processing device 410. The second processing device 420 is a separate device to the first processing device 410, with the two being in communicably connected via a network connection. The second processing device 420 may be a games console remote to the first processing device 410, for example, or a cloud-based processing arrangement (cloud gaming server) which is configured to execute an instance of a video game. The display device 430 configured to display images generated by the first processing device 410 to both the first and the second user, with the displayed images being dependent upon both the first and second instances of the video game. Turning to Figure 5, this Figure schematically illustrates a configuration of the first processing device 410 of Figure 4 in more detail. The processing device 410 comprises a processor 500, a communication unit 510, an input control unit 520, and an image generation unit 530. These functions may be implemented by the CPU 20, GPU 30, and data port 60 of the entertainment system 10 of Figure 1, for example; however any suitable processing hardware may be used to realise this functionality. The processor 500 is configured to execute a first instance of the video game responsive to inputs received from the first control device 400, such as those to control a user's avatar within the game environment. These inputs may include those which enable a second user to join the game via the second instance, such as issuing an invitation to that second user or configuring the first instance to enable other users to join. The processor 500 may be configured to adapt one or more settings of the first instance of the video game in dependence upon one or more parameters (such as video settings) associated with the second instance of the video game, the output video of the second instance of the video game (should this be provided by the second processing device 420), and / or properties of the network connection between the first processing device 410 and the second processing device 420. This can enable the presentation of the first instance of the video game to be in keeping with (that is, appearing similar or the same) that of the second instance (or the expected presentation of the second instance, in the case that it is not displayed); this can be particularly desirable in an implementation in which both instances are displayed 7 simultaneously in a split-screen fashion. In the case in which only the first instance is displayed, this may still provide advantages in providing video content that accounts for the display settings of the second instance such as brightness which may be important considerations for ensuring user comfort and content visibility. The processor 500 may also, or instead, be configured to modify a camera viewpoint associated with the first instance of the video game in dependence upon an output by the second instance of the video game. For instance, based upon data indicating the location of the second user's avatar in the second instance of the video game a camera viewpoint may be adjusted to ensure that both user's avatars are visible in the same image generated from the first instance of the video game. This can aid an implementation in which a single image is displayed to the users which is representative of the gameplay of both users. In some cases, it is considered that based upon such data the output video may be switched between split-screen and a single image in dependence upon a threshold distance between the first and second user's avatars such that when the threshold distance is exceeded the display is changed to a split-screen view. In some implementations the processor 500 may be configured to identify a latency associated with the receiving of data from the second processing device (such as a latency associated with the network connection and / or a processing time in transmitting inputs from the first processing device 410 to the second processing device 420), and to apply an input latency and / or display latency to the execution of the first instance in dependence upon this. In other words, the processing of the first instance can be adapted to provide an equal (or at least similar) latency to that of the second instance so that each user is able to interact with their respective game instances in a mutually consistent manner. An input latency refers to delaying the provision of the inputs to the game, whilst a display latency refers to delaying the display of images of the game to the user. The introduced latency may be a fixed value which is representative of an average or expected latency, or it may be responsive to live measurements of said latency. The processor 500 may be further configured to adapt settings associated with the first instance of the video game and / or instruct the adapting of settings of the second instance of the video game so as to manage a local processing load or the like. For instance, some hardware may find that executing a game while decoding received video content represents a significant processing burden - in such a case, the video quality of either instance (or indeed both instances) may be modified so as to reduce this burden and ensure that processing can be effectively managed. The communication unit 510 is configured to receive data from a second processing device 420 via a network, the data corresponding to a second instance of the video game being executed concurrently with the first instance. The communication unit 510 may be further configured to perform other communications, such as the transmission of inputs described with reference to the input control unit 520. In some implementations, the communication unit 510 is configured to transmit information to the second processing device 420 comprising information about the initiation of a game session - such as a location of the first user in the in-game environment, or other game state information such as a current stage, user loadout, and quest. The data corresponding to the second instance of the video game may be in any suitable format. In some implementations, the communication unit 510 is configured to receive data comprising output video of the second instance of the video game (optionally with the associated audio). Alternatively, or in addition, the communication unit 510 is configured to receive data comprising the results of one or more simulations (such as physics simulations for in-game interactions by the second user) performed by the second instance of the video game, wherein the results of the one or more simulations are provided to the first instance of the video game. The communication unit 510 can be considered optional, as a number of different implementations may not require any communication - such as when the two game instances are being executed locally by the processor 500 of the first processing device 410. In such a case, the two game instances may be configured to communicate directly without the need of a separate communication unit as is the case when the second instance is being executed by a remote server or the like. In such implementations, the functionality of the second processing device is realised by the first processing device as appropriate to provide two separate instances of the same game using a single device. The input control unit 520 is configured to provide inputs received from the first control device to the first instance of the video game, and to transmit inputs received from the second control device to the second processing device via the network connection. This may be performed by an in-game function associated with the first instance of the video game (or a separate game-specific tool which is executed alongside the video game), or it may be handled externally to the game such as by a system-level function provided by an operating system run by the first processing device 410. The image generation unit 530 is configured to generate images for display in dependence upon both the first and second instances of the video game, with these images being provided for output to both the first and second user by the display device 430 of Figure 4. In some implementations, the image generation unit 530 is configured to generate a split-screen image comprising output video of each of the first and second instances of the video game. However, this is not considered to be limiting; in some cases it may be preferable that the first instance of the video is used to generate images for display which are representative of the gameplay of both users. It is also envisaged that the format is a dynamic one which is responsive to in-game events or conditions - such as based upon user proximity in the game environment such that as the users move apart a split-screen is preferred, or switching to a single screen during cut-scenes or the like. As described above, in the exemplary implementations discussed each of the first instance and second instance of the video game are capable of supporting a single player only. However, it is considered that the same techniques may be extended to any case in which the number of users exceeds the number of users that a single instance of a video game is capable of supporting. Similarly, the number of processing devices is not limited to two, but could be increased to any suitable number. While discussed above with the first instance of the video game being executed locally, it is also considered that the first processing device could be implemented as a thin client or the like which decodes video received from two remote game instances. This may be particularly suitable for low-powered devices such as mobile phones or portable gaming consoles. This thin client may comprise the communication unit 510, input control unit 520, and image generation unit 530 whilst the functionality of the processor 500 is provided remotely (such as by a games console or cloud gaming server). As such, the thin client is configured to receive inputs from the two or more control devices, route these to the appropriate game instances, and receive video which is to be displayed to the users who are local to the thin client. The arrangement of Figure 4 is an example of a system for executing a video game which comprises a first processing device configured to receive inputs from each of two or more control devices associated with respective users, comprising a first control device associated with a first user and a second control device associated with a second user, the first processing device being communicably connected with a second processing device. Figure 5, which illustrates the first processing device, is an example of an arrangement which can be implemented using a processor (for example, a GPU and / or CPU located in a games console or any other computing device) that is operable to receive inputs from each of the two or more control devices to control respective instances of a video game, and in particular is operable to: execute a first instance of the video game at the first processing device; receive data from a second processing device via a network, the data corresponding to a second instance of the video game being executed concurrently with the first instance; provide inputs received from the first control device to the first instance of the video game; transmit inputs received from the second control device to the second processing device; generate images for display in dependence upon both the first and second instances of the video game; and display the generated images to both the first and the second user via the same display device, wherein the first instance of the video game is responsive to inputs received from the first control device, and the second instance of the video game is responsive to inputs received from the second control device. This functionality may be provided in accordance with any of the hardware configurations described elsewhere in this disclosure; for instance, the first processing device may be implemented as the entertainment system 10 of Figure 1. In this case, processing is performed by the CPU 20 and / or GPU 30. Figure 6 schematically illustrates a method for executing a video game by a first processing device configured to receive inputs from each of two or more control devices associated with respective users, comprising a first control device associated with a first user and a second control device associated with a second user. A first instance of the video game is responsive to inputs received from the first control device, and a second instance of the video game is responsive to inputs received from the second control device. A step 600 comprises executing a first instance of the video game at the first processing device. A step 610 comprises receiving data from a second processing device via a network, the data corresponding to a second instance of the video game being executed concurrently with the first instance. A step 620 comprises managing inputs received from the two or more control devices, with the management comprising providing inputs received from the first control device to the first instance of the video game and transmitting inputs received from the second control device to the second processing device. A step 630 comprises generating images for display in dependence upon both the first and second instances of the video game. A step 640 comprises displaying the generated images to both the first and the second user via the same display device. When users interact with applications in accordance with the above implementations, it is considered that the users will have an overlap in the respective videos output by their respective application instances. In other words, it is considered that at least some of the time the videos displayed for each user would be similar or at least share a number of common elements, as would their corresponding audio. An example of this is when playing a game - each of the users may proximate to one another in the game environment, and therefore would be provided with similar video / audio. Similarly, when viewing free-viewpoint media content the same scenarios may arise - such as if two users sit next to each other in an immersive sports stadium experience. Even in the case in which users in an environment aren't close to one another there may be a number of shared audio elements, such as background music or global sound effects (such as announcements). Audio elements here refer to component parts of the audio, such as sounds associated with a particular sound source. Reproducing such audio in parallel may be distracting to a user in some cases if shared audio reproduction hardware is used (such as the speakers associated with a display, rather than using individual headsets or the like), especially if there is a latency between the streams as this can cause the same audio to be presented with a temporal offset. While issues resulting from the reproduction of parallel audio streams can be circumvented by using separate audio reproduction devices (such as each user having a respective pair of headphones), this causes users to be more isolated in the shared environment as they are less able to hear real-world sounds. It is also considered that the rendering and transmission of duplicated content can be inefficient, placing an unnecessary burden on the system and network. These negative effects are amplified as the number of users increases; therefore while the below discussion primarily focuses upon an implementation in which only two users are interacting, it is considered that it would be desirable to extend these teachings to arrangements directed toward a greater number of users. It is therefore desirable to exploit the similarities between the respective audio outputs where possible to improve efficiency of the system. Figure 7 schematically illustrates a method which seeks to achieve this aim. The steps of Figure 7 may be performed by any suitable processing hardware in the system; in some cases, the local processing device may perform the audio modification, while in others it may be preferable for such processing to be performed by a server (such as one which executes an instance of the application). The audio modification processing may be performed by the first or second instance of the application, for example, or by a standalone process which operates independently of the execution of the instances of the application. As such, any specific discussion of transmission of information and location of processing is omitted in the below discussion with the skilled person being able to determine these implementation details in accordance with the details of a given implementation. A step 700 comprises executing a first and second instance of an application. These instances may be executed at any suitable devices in accordance with the above discussion; for instance, one may be executed at a local processing device with another being executed remotely (such as at a second processing device associated with a second user, or at a server), or both may be executed remotely. The implementation of this method is able to be adapted to account for these different options without undue burden upon the skilled person. A step 710 comprises obtaining audio output data from each of the first and second instances of the application. The audio output may be obtained from a video stream which includes both visual and audio content, for instance, or the audio may be obtained from a separate stream to any video content. The audio output data may comprise the audio output itself; alternatively, or in addition, other data which characterises the audio output or an application state corresponding to the audio output may be obtained. For instance, information about the location and orientation of an in-game camera / microphone may be obtained, or information indicating the start / end of a cutscene. Information representing common elements between the instance's respective audio outputs may also be obtained, such as identifying sound effects which are common between the instances - for instance, background music or global announcements. Such data may be obtained separately to the audio, or it may be encoded as metadata alongside the audio, for example. A step 720 comprises identifying an overlap between the respective audio outputs. Here, an overlap can be an overlap in content (such that the audio outputs comprise the same audio elements), or an overlap in reproduction (such that the audio outputs comprise different audio elements being provided at the same time). In some cases, step 720 may comprise an identification of both of these overlap types rather than being limited to a single one. This step may utilise any suitable processing or data to allow overlaps to be identified, with the audio itself and / or information about the audio being used as the basis of the processing or the data source. In some cases it is necessary to consider latency between the application instances; this can be a processing / transmission latency, or it may be a latency which arises due to a simulated speed of sound within a virtual environment in some applications (in other words, an apparent audio latency may be identified due to each user in an environment being a different distance from the sound source). In this case, the comparison may be tailored to the identified latency to ensure that the corresponding parts of the audio are being compared. Of course, this latency may be vary for different audio elements within the content; this may also be factored into the analysis where appropriate. A first example of the processing is that of processing the respective audio outputs directly to identify overlaps; for instance, samples of the audio outputs may be processed to extract features (such as a frequency analysis) which are then compared. A subtractive approach could also be taken, in which a sample of one instance is subtracted from a corresponding sample of the other instance to determine a residual. Similarly, a sound recognition process may be performed on the audio so as to identify common elements. This can be performed in a number of different ways. For instance, in a first case this may comprise searching both audio outputs for audio elements that would be known to be shared such as global announcements presented to all users independent of location. Alternatively, or in addition, a sound recognition process can be performed on the audio output of one of the instances with the other of the audio outputs being searched for corresponding audio. In some implementations, the application instances may output information about their respective audio outputs directly. This may be a pre-processed representation of the audio output to make a comparison more efficient, for example, or may include semantic information which describes or directly identifies elements within the audio output. For instance, this could include the output of filenames (and optionally timing information), or at least a type of sound (as different users may have assigned different sounds to the same event - such as having selected different commentator voices or having their application provide announcements in different languages). Information output by the application instances can also include information about the locations of a virtual microphone in a virtual environment, and / or information about the locations of virtual sound sources. Information about these can be used to identify an overlap between the audio in each instance. For instance, if two users are standing side-by-side in a virtual environment then it can be assumed that the audio outputs are largely identical. Based upon relative locations, it can be determined whether one user would hear something that the other cannot, or relative sound levels can be determined, based upon a sound propagation model (or at least a rough estimation, for improved efficiency) associated with a given virtual environment. Should a sound source location be above a threshold distance from a user's position in a virtual environment, it can be assumed that the sound is either global (such as an announcement) or sufficiently loud to be heard by both users - and therefore an overlap can be identified on this basis. Metadata may also be provided with (or encoded as a part of) the audio output which categorises one or more of the sounds present in the audio output. This can be useful for sounds which do not have a particular location in a virtual environment associated with an application, for instance, with categories indicating whether a sound is 'general' (that is, the sounds of the environment), 'user-specific' (such as audio effects indicating a user having low health), or 'global' (such as background music or announcements). Categories can of course be selected freely; for instance, 'general' can be subdivided into 'near', 'medium distance', and 'distant', or more / fewer categories may be utilised. Another example of data that can be utilised is that of information indicating the presence of subtitles, either found in the metadata associated with a video stream or identified based upon image processing of one or more image frames in the case of hard-coded subtitles. The presence of subtitles may be indicative of a cutscene, or at least of important audio that should be afforded a high level of priority should a selection of audio elements to reproduce be made. Cut scenes may also be identified separately, such as through metadata output alongside the video content or from the video content itself (such as through a watermark added to indicate the start of a cutscene, or an identification that the instances are outputting identical video content). A step 730 comprises modifying the audio that is to be presented to the users at the local processing device so as to generate a combined audio output. This can be performed in a number of ways, each of which may be utilised in combination where appropriate. In general, this step may comprise one or more of modifying the audio output of one or more of the audio instances, mixing those outputs, or causing one or more of the application instances to generate a different audio output. Modifying the audio output of one or more of the audio instances may include processing to remove or mute common audio elements amongst the audio outputs, for instance, or audio elements which are otherwise not to be reproduced in the combined audio output. Mixing the audio outputs may include combining the audio outputs in a manner which varies the contribution of one or more audio elements of the component audio outputs. This can include reducing the volume of one or more audio elements in the mix, for example, or even discarding an audio output associated with one of the instances. The mixing may also comprise an interpolation of one or more of the audio elements where suitable. For example, the sounds themselves may interpolated so as to generate a representation which is indicative of sounds associated with each of the instances, or an interpolation may be performed so as to change an apparent location of a sound source relative to a listener. For instance, an interpolation may be performed so as to generate a combined audio output having an apparent location which is between the locations (such as at a midpoint), in a virtual environment of the application, associated with the respective audio outputs. This can ease a feeling of discomfort that may arise if sound effects are reproduced with an apparent location relative to the listener which is too far removed from what would be expected. Causing one or more of the application instances to generate a different audio output comprises providing an instruction (or information upon which an instruction is generated by the application instance) to the application instance to modify its operation in respect of audio generation. This can include a case in which audio output is terminated for an application, or the audio output is a stream 15 comprising no data (in the case that this aids compatibility with particular streaming formats, for instance). Similarly, particular elements can be caused to be omitted from the audio output; this may be particular sounds or classes of sounds (such as 'global announcements'), for example. Those omitted audio elements may be those which would appear in both audio outputs, or those which cause a clash with audio elements in those other audio outputs. In some implementations, the modification step may be dependent upon a prioritisation system to determine which audio elements should be reproduced in the case of a clash. A clash here is considered to be any combination of reproduced audio elements which is undesirable - such as one audio element obscuring another (such as an explosion during dialogue), incompatibility between the audio elements (such as two different sets of background music), or the audio elements being considered to be too distinct from one another (such as if one instance is providing 'fun' sounds while another is providing 'scary' sounds). By assigning a priority value to particular audio elements and / or classes of audio elements, such conflicts can be resolved in an efficient manner. In the case that a conflict is identified between sounds of equal priority, user preferences may be used to determine how to proceed, or a particular instance of the application may be designated as the 'primary instance' to which other instances defer. User preferences may be predefined, such as a user profile indicating which sounds or types of sounds should be prioritised, or this may be performed live. In some instances, the predefined preferences are used as the basis for audio reproduction but can be modified on-the-fly by users. These preferences can be defined with any suitable degree of granularity - such as particular sounds or applications - and can be defined in respect of particular combinations of content / applications / users as desired. In some instances, the user or users may be presented with a UI element comprising a cross-fader style functionality to enable a fine-tuning of the combined audio output of the instances of the application. Finally, a step 740 comprises reproducing the audio at the local processing device; this may be by a single display device (and associated audio reproduction elements, such as a surround sound system), or the reproduction may be divided amongst a number of devices. For instance, in the case that multiple displays are provided it may be the case that audio elements associated with the first instance and both instances are reproduced at one display device, with the second reproducing only those elements associated with the second instance. In the case that the audio reproduction means comprises a directional audio output, each of the users may be targeted with their corresponding instance's unique audio while the shared audio elements are played via non-directional audio reproduction means. Implementations of a method in accordance with that described above therefore provide the ability to provide two distinct gameplay instances with their respective accompanying audio via a single sound stage. Figure 8 schematically illustrates a system for providing a synchronised multi-user application experience at a local processing device. The system comprises a first application processing unit 800, a second application processing unit 810, an audio analysis unit 820, a mixing unit 830, and an audio reproduction unit 840. Not shown in this Figure are the local processing device and any remote processing devices or servers; this is because the functional units shown may be distributed between such devices as appropriate for a given implementation. The first application processing unit 800 is configured to execute a first instance of the application responsive to inputs received from the user of a first input device associated with the local processing device. The first application processing unit 800 may be implemented by any suitable arrangement of hardware, such as a CPU and GPU in communication with one another, as may the second application processing unit 810. The second application processing unit 810 is configured to execute a second instance of the application responsive to inputs received from the user of a second input device associated with the local processing device, wherein at least one of the first and second application processing units is remote to the local processing device. In some implementations, the first application processing unit 800 is located at the local processing device and the second application processing unit 810 is located at a remote processing device or a server. Alternatively, the first application processing unit 800 is located at a remote processing device or a server, and the second application processing unit 810 is also located at a remote processing device or a server. In the latter case, the functionality of the first and second application processing units may be realised by the same server. The key feature regarding the first and second application processing units is that they execute separate instances of the same application in a multi-user configuration; as such, the specific locations of the instances are able to be selected freely while the inputs to both are still received by the local processing device shared by the users. The audio analysis unit 820 is configured to analyse audio outputs associated with each of the instances of the application to identify conflicting audio elements amongst the audio elements associated with each audio output. Conflicting audio elements are those which, when both reproduced, would be considered inefficient due to duplication, would cause auditory discomfort for a listener, or would otherwise impair the listening experience. For instance, the audio analysis unit 820 may be configured to identify an audio element which is present in both audio outputs as a conflicting audio element; alternatively, or in addition, the audio analysis unit is configured to identify an audio element in an instance of the application which would impair the audibility of an audio element in the other instance of the application as a conflicting audio element. Examples of such conflicting audio are discussed above, and may include duplicated background music and loud noises during dialogue, for instance. The audio analysis unit 820 may additionally be configured to analyse data output by the respective instances of the application to identify conflicting audio elements, the data being indicative of an application state of the corresponding instance. This may include information such as the location of a virtual microphone in respective virtual environments (where a similar location between the instances would be indicative of a significant overlap in audio elements in the corresponding audio outputs), or information about what is happening in the application - such as the occurrence of an event (which may correspond to a specific audio element) or the start of a cutscene in a game, for example. Similarly, the audio analysis unit 820 may be configured to analyse video data output by the respective instances of the application to identify conflicting audio elements, the video data being indicative of the content of the audio outputs. In some implementations this may comprise processing the video directly to identify particular elements or events from which information about an audio output can be derived; alternatively, or in addition, this approach may include identifying watermarks or the like in the video content which are inserted to indicate the presence of one or more audio elements or events (such as cutscenes). For instance, the analysis of video may include the identification of events within the content. An example of this is identifying that the score has changed in a sports game (for instance, from a change on the scoreboard), with the presence of a corresponding audio element of a score announcement being inferred from this. The analysis of video may also perform a comparison between the respective views being presented in each instance of the application; if the views are similar then it can be assumed that the users are proximate to one another in a virtual environment and therefore they would experience similar audio. It can also be considered that if both users are presented with a view of the same element then they will both be presented with audio corresponding to that element - and as such duplication of audio can be expected. In order to implement this process, it may be preferred to use a machine vision process which has been trained on the specific application (or a group of similar applications) to obtain improved results; this training may be based upon pairs of assets and associated audio elements, for instance. The audio analysis unit 820 may be configured to obtain information indicating class and or identification information for one or more audio elements within an audio output. This may be based upon information encoded as a part of the audio or metadata associated with the audio and / or video output of the application, for example; an application may be configured to output such information alongside the audio-visual output. In some cases the identification of an audio element may be based upon processing of the audio and / or visual content output by the application, for instance using a trained machine learning model which is trained on that application. Once identified, class information can be derived based upon a locally stored look-up table or the like. Classes of audio elements refers to a type of audio element, typically defined in dependence upon how the audio is perceived by a listener or how widely-heard the audio is. For example, classes may include 'background music' or 'global sounds', 'near' or 'far' sounds (for instance, based upon typical volume), and / or 'user-specific' or 'team-specific' sounds. A single audio element may be associated with multiple classes where appropriate. The audio analysis unit 820 may be configured to obtain information indicating a relative latency between the first and second instances of the application, and to identify conflicting audio in dependence upon the relative latency. This may be advantageous in that it enables conflicts to be identified more readily - if there is a latency between the instances then the audio elements may be reproduced at different times, thereby meaning that duplication of sounds may not be identified if latency is not considered. Information about the latency is used to ensure that corresponding times in the respective audio outputs are being compared; for instance, audio analysis results at time't' for one instance can be compared against the results for time 't+latency' in the other instance. Similarly, the relative locations of users within a virtual environment of an application may be considered - given the relatively low speed of sound, it may be considered that a latency is introduced due to different listener locations with respect to a sound source. Such a latency may be determined based upon information about the spatial arrangement of elements in the application, which may be output by one or both of the application instances, or such latency information may be generated by one or both of the application instances themselves and output to the audio analysis unit 820. The audio analysis unit 820 may be configured to operate on the respective audio output streams in any suitable manner. In one example, a circular buffer is maintained for each instance which comprises a most recent portion of the audio output of the corresponding instance. The size of this buffer may be determined freely, although typically a small buffer is used to reduce storage requirements-storing half a second of the audio output of each instance may be sufficient, for instance. The contents of these buffers can be compared in a continuous manner to identify any conflicting audio. The audio analysis unit 820 is therefore configured to analyse the audio outputs of one or both of the instances of the application, and optionally video or application data output by one or both of the instances. This is for the purpose of identifying audio conflicts between the instances - that is, audio which is duplicated, or which is incompatible (that is, the reproduction of an audio element from one instance would impair the user's listening to audio elements associated with the other instance). The mixing unit 830 is configured to generate combined audio representative of the respective audio outputs associated with each of the instances of the application, wherein one or more conflicting audio elements are omitted from the combined audio. In some cases this omission may be a lowering of the volume of the audio element to a significantly lower level, such that a listener is less able to perceive that audio element in the combined audio. The mixing unit 830 may be configured to modify the audio outputs directly to obtain a desired combined audio; however, in some cases it may be preferred that the mixing unit 830 is configured to cause one or both of the instances of the application to modify their audio outputs in response to identifying conflicting audio elements. This can be by generating an instruction to one or both instances as appropriate to indicate particular audio elements or classes of audio elements that should be omitted from an audio output - or instructing an instance to cease audio output altogether. The mixing unit 830 may be configured to perform an interpolation process such that the apparent sound source location associated with a conflicting audio element is changed in the combined audio to be a location between the locations of the audio element in the respective audio outputs of the first and second instances of the application. While this can mean that an audio element is not presented at the correct location for either listener, this can reduce the occurrence of extreme differences between an expected and actual sound source location for the user of an application instance for which the output of an audio element is to be terminated. This interpolation can be performed directly upon the audio should the audio output comprise spatial information or use a three-dimensional audio format; alternatively, an instance of the application can be caused to modify its generation of audio to reflect the interpolated location. In determining how to handle conflicting audio elements, the mixing unit 830 may be configured to generate combined audio in dependence upon a priority associated with an audio element and / or class of audio elements. This can assist in resolving conflicts, as the use of a priority value can indicate which of the conflicting audio elements should be retained. Alternatively, or in the case that priorities are equal, a particular instance of the application may be regarded as the primary instance for audio purposes - with all audio elements from this instance having priority over those from other instances. The audio reproduction unit 840 is configured to output the combined audio; that is, to reproduce the audio for listening by the users of the local processing device. This may be performed alongside the display of corresponding video content on one or more display devices associated with the local processing device. The reproduction of audio may be performed utilising any suitable arrangement of hardware, such as a surround sound system associated with one or more displays or integrated speakers provided as a part of a display device. The arrangement of Figure 8 is an example of a processor (for example, a GPU, TPU, and / or CPU located in a games console or any other computing device) that is operable to provide a synchronised multi-user application experience at a local processing device, and in particular is operable to: execute a first instance of the application responsive to inputs received from the user of a first input device associated with the local processing device; execute a second instance of the application responsive to inputs received from the user of a second input device associated with the local processing device, wherein at least one of the first and second application processing units is remote to the local processing device; analyse audio outputs associated with each of the instances of the application to identify conflicting audio elements amongst the audio elements associated with each audio output; generate combined audio representative of the respective audio outputs associated with each of the instances of the application, wherein one or more conflicting audio elements are omitted or changed in the combined audio; and output the combined audio. Figure 9 schematically illustrates a method for providing a synchronised multi-user application experience at a local processing device. This method may be implemented in accordance with the features described with reference to the system Figure 8, or any of the other features described above. A step 900 comprises executing a first instance of the application responsive to inputs received from the user of a first input device associated with the local processing device. A step 910 comprises executing a second instance of the application responsive to inputs received from the user of a second input device associated with the local processing device, wherein at least one of the first and second application instances is executed remote to the local processing device. A step 920 comprises analysing audio outputs associated with each of the instances of the application to identify conflicting audio elements amongst the audio elements associated with each audio output. A step 930 comprises generating combined audio representative of the respective audio outputs associated with each of the instances of the application, wherein one or more conflicting audio elements are omitted from the combined audio. A step 940 comprises outputting the combined audio. 5 The techniques described above may be implemented in hardware, software or combinations of the two. In the case that a software-controlled data processing apparatus is employed to implement one or more features of the embodiments, it will be appreciated that such software, and a storage or transmission medium such as a non-transitory machine-readable storage medium by which such software is provided, are also considered as embodiments of the disclosure. 10 Thus, the foregoing discussion discloses and describes merely exemplary embodiments of the present invention. As will be understood by those skilled in the art, the present invention may be embodied in other specific forms without departing from the spirit or essential characteristics thereof. Accordingly, the disclosure of the present invention is intended to be illustrative, but not limiting of the scope of the invention, as well as other claims. The disclosure, including any readily discernible variants of the 15 teachings herein, defines, in part, the scope of the foregoing claim terminology such that no inventive subject matter is dedicated to the public.

Claims

1. A system for providing a synchronised multi-user application experience at a local processing device, the system comprising:a first application processing unit configured to execute a first instance of the application responsive to inputs received from the user of a first input device associated with the local processing device;a second application processing unit configured to execute a second instance of the application responsive to inputs received from the user of a second input device associated with the local processing device, wherein at least one of the first and second application processing units is remote to the local processing device;an audio analysis unit configured to analyse audio outputs associated with each of the instances of the application to identify conflicting audio elements amongst the audio elements associated with each audio output;a mixing unit configured to generate combined audio representative of the respective audio outputs associated with each of the instances of the application, wherein one or more conflicting audio elements are omitted or changed in the combined audio; andan audio reproduction unit configured to output the combined audio.

2. A system according to claim 1, wherein the first application processing unit is located at the local processing device and the second application processing unit is located at a remote processing device or a server.

3. A system according to claim 1, wherein the first application processing unit is located at a remote processing device or a server, and the second application processing unit is located at a remote processing device or a server.

4. A system according to any preceding claim, wherein the audio analysis unit is configured to identify an audio element which is present in both audio outputs as a conflicting audio element.

5. A system according to any preceding claim, wherein the audio analysis unit is configured to identify an audio element in an instance of the application which would impair the audibility of an audio element in the other instance of the application as a conflicting audio element.

6. A system according to any preceding claim, wherein the audio analysis unit is configured to analyse data output by the respective instances of the application to identify conflicting audio elements, the data being indicative of an application state of the corresponding instance.

7. A system according to any preceding claim, wherein the audio analysis unit is configured to analyse video data output by the respective instances of the application to identify conflicting audio elements, the video data being indicative of the content of the audio outputs.

8. A system according to any preceding claim, wherein the audio analysis unit is configured to obtain information indicating class and or identification information for one or more audio elements within an audio output.

9. A system according to any preceding claim, wherein the audio analysis unit is configured to obtain information indicating a relative latency between the first and second instances of the application, and to identify conflicting audio in dependence upon the relative latency.

10. A system according to any preceding claim, wherein the mixing unit is configured to cause one or both of the instances of the application to modify their audio outputs in response to identifying conflicting audio elements.

11. A system according to any preceding claim, wherein the mixing unit is configured to perform an interpolation process such that the apparent sound source location associated with a conflicting audio element is changed in the combined audio to be a location between the locations of the audio element in the respective audio outputs of the first and second instances of the application.

12. A system according to any preceding claim, wherein the mixing unit is configured to generate combined audio in dependence upon a priority associated with an audio element and / or class of audio elements.

13. A method for providing a synchronised multi-user application experience at a local processing device, the method comprising:executing a first instance of the application responsive to inputs received from the user of a first input device associated with the local processing device;executing a second instance of the application responsive to inputs received from the user of a second input device associated with the local processing device, wherein at least one of the first and second application instances is executed remote to the local processing device;analysing audio outputs associated with each of the instances of the application to identify conflicting audio elements amongst the audio elements associated with each audio output;generating combined audio representative of the respective audio outputs associated with each of the instances of the application, wherein one or more conflicting audio elements are omitted from the combined audio; andoutputting the combined audio.

14. Computer software comprising instructions which, when the software is executed by a computer, causes the computer to carry out the method of claim 13.

15. A non-transitory machine-readable storage medium which stores computer software according to claim 14.

Citation Information

Patent Citations

  • Enabling local split-screen multiplayer experiences using remote multiplayer game support

    US20230381640A1