Welcome!

Blog Feed Post

Lync 2013 Persistent Chat HA\DR Deep Dive Pt. 2

In part 1 of this blog we discussed the architecture and design of Lync 2013 pChat HA\DR components. Now we will discuss how this design behaves with various failures within a Lync infrastructure. These different failure scenarios are based off our Disaster Recovery diagram from part 1 (Figure 1). We will cover the following failure scenarios:

  1. Lync Front End Pool failure
  2. Complete Site Failure
  3. pChat Pool failure
  4. Site Recovery

 

Scenario 1A: Lync FE Pool failure (Figure 1)

Figure 1: EEPool1 fails in East Datacenter

In this scenario our Front End Pool (EEPool1) in the East Datacenter has a complete failure. In order for users to get reconnected to pChat1 we will need to failover our paired pool from the East Datacenter to the West Datacenter. This will allow EEPool2 to route pChat connections to our pChat pool. This scenario would have the same failover steps regardless of if all pChat servers are active in oncdatacenter or split between the two (discussed in part 1).

  1. Failover the pool from EEPool1 to EEPool2
    Invoke-CsPoolFailOver –PoolFqdn “eepool1.domain.com” –DisasterMode2

1 The same steps will be performed regardless of whether all of the pChat servers are active in a single datacenter or they are active in both datacenters.

2 The –DisasterMode parameter is used with the Invoke-CsPoolFailOver cmdlet because our Pool has completely failed.

Once the failover is complete users can reconnect and all pChat functionality will be restored.

 

Scenario 1B: Lync FE Pool failback

Once our Front End Pool (EEPool1) in the East Datacenter comes back online and all services are restored we need to perform a failback.

  1. Invoke-CsPoolFailback -PoolFQDN "eepool1.domain.com" –Verbose

 

Scenario 2A: Complete Site failure (Figure 2)

Figure 2: Entire East Datacenter failure – all services in this Datacenter are down (FE, pChat, SQL, File Store, Edge)

Next, we look at a complete Lync site failure. This encompasses all Lync services including FE, pChat, SQL, File Store, and Edge. In order to restore all User services, including pChat, we will need to activate all these services in the secondary datacenter. In order to shorten this blog a little there are links to the edge failover process at end.

  1. Failover Pool from EEPool1 to EEPool2
    1. Invoke-CsPoolFailOver –PoolFqdn “eepool1.domain.com” –DisasterMode2
  2. Prepare pChat Backup DB, Commit TLogs, and bring DB online
    1. Remove log shipping from the Persistent Chat Server Backup Log Shipping database
      1. Using SQL Server Management Studio, connect to the database instance where the Persistent Chat Server backup MGC database is located
      2. Open a query window to the master database
      3. Drop Log shipping  - exec sp_delete_log_shipping_secondary_database mgc
    2. Copy any uncopied backup files from the backup share to the copy destination folder of the backup server.
    3. Apply any unapplied transaction log backups in sequence to the secondary database
      1. "How to: Apply a Transaction Log Backup (Transact-SQL)" at http://go.microsoft.com/fwlink/p/?linkid=247428
    4. Bring the backup MGC database online. Using the query window that was opened in step 2a, end all connections to the MGC database, if there are any:
      1. \exec sp_who2 to identify connections to the mgc database.
      2. \kill <spid> to end these connections.
    5. Bring the database online
      1. \restore database mgc with recoverSet pChat Pool state as failed over - after this is completed the MGC backup DB will now serve as the primary database.
  3. Set pChat Pool state as failed over - after this is completed the MGC backup DB will now serve as the primary database.
    1. Set-CsPersistentChatState -Identity “PersistentChatServer:pchatpool1.domain.com” –PoolState FailedOver
    2. Use the Get-CSPersistentChatState to verify it is marked as “Failed Over”
  4. Set the pChat Active servers in the secondary datacenter (West datacenter in our example) – at this point we want to make sure all the pChat servers in the secondary datacenter are active
    1. Set-CsPersistentChatActiveServer –Identity “global” –ActiveServers @{Add="pchatserver6.domain.com"}1
  5. (Optional but recommended) – use the Install-CsMirrorDatabase cmdlet to establish a High Availability mirror for the backup database that now serves as the primary database.

1 Add all pChat servers (using FQDN) in secondary datacenter (West datacenter) separated by commas in between quoted servers names.

Once the failover is complete users can reconnect and all pChat functionality will be restored using services from the secondary (paired) datacenter.

 

 Scenario 2B: Complete Site failback

 Once the East datacenter comes back online and all services come online we need to perform a failback.

  1. First lets failback our pool to the Primary datacenter (East Datacenter)
    1. Invoke-CsPoolFailback -PoolFqdn “eepool1.domain.com”
  2. Clear all servers from the Persistent Chat Server Active Server list
    1. Set-CsPersistentChatActiveServer (no switches)
  3. If you enabled DB mirroring on the backup DB in the secondary datacenter disable mirroring
    1. Using SQL Server Management Studio, connect to the backup MGC instance.
    2. Right-click the MGC database, point to Tasks, and then select Mirror.
    3. Click Remove Mirroring, and then select OK.
  4. Back up the MGC database so that it can be restored to the new primary database
    1. Using SQL Server Management Studio, connect to the backup MGC instance.
    2. Right-click the MGC database, point to Tasks, and then click Back Up. The Back up Database dialog box appears.
    3. In Backup type, select Full.
    4. For Backup component, click Database.
    5. Either accept the default backup set name suggested in Name, or enter a different name for the backup set.
    6. <Optional> In Description, enter a description of the backup set.
    7. Remove the default backup location from the destination list.
    8. Add a file to the list by using the path to the share location that you established for log shipping. This path is available to the primary database and to the backup database.
    9. Click OK to close the dialog box and begin the backup process.
  5. Restore the primary database by using the backup database created in the previous step.
    1. Using SQL Server Management Studio, connect to the primary MGC instance.
    2. Right-click the MGC database, point to Tasks, point to Restore, and then click Database. The Restore Database dialog box appears.
    3. Select From Device.
    4. Click the browse button, which opens the Specify Backup dialog box. In Backup media, select File. Click Add, select the backup file that you created in step 3, and then click OK.
    5. In Select the backup sets to restore, select the backup.
    6. Click Options in the Select a page pane.
    7. In Restore options, select Overwrite the existing database.
    8. In Recovery State, select Leave the database ready to use.
    9. Click OK to begin the restoration process.
  6. Set pChat Pool state as Normal -
    1. Set-CsPersistentChatState –Identity “PersistentChatServer:pchatpool1.domain.com” –PoolState Normal
    2. Use the Get-CSPersistentChatState to verify it is marked as “Normal”
  7. Set the pChat Active servers to what they were prior to failover.
    1. Set-CsPersistentChatActiveServer –Identity “global” –ActiveServers @{Add=”pchatserver6.domain.com”}1
  8. pChat clients should now reconnect to the Active pChat servers
  9. Setup Log shipping again to the secondary datacenter for DR purposes
    1. Follow steps here - http://technet.microsoft.com/en-us/library/jj204653.aspx

1 Add all pChat servers (using FQDN) that were active before failover separated by commas

 

Scenario 3A: pChat Pool Server failure (Figure 3)

Figure 3: pChat Pool failure in East Datacenter

The third failure scenario that we will explore is a pChat Pool Server failure in which all active pChat servers are located in the East Datacenter1. We will need to failover all pChat services to the secondary datacenter. Once we complete the failover of the pChat services, EEPool1 will route pChat traffic to the pChat servers located in the secondary datacenter.

  1. Prepare pChat Backup DB, Commit TLogs, and bring DB online – in order to accomplish these tasks follow step 2 from Scenario 2A from the complete site failure above.
  2. Set pChat Pool state as failed over – after this is completed the MGC backup DB will now serve as the primary database.
    1. Set-CsPersistentChatState –Identity “PersistentChatServer:pchatpool1.domain.com” –PoolState FailedOver
    2. Use the Get-CSPersistentChatState to verify it is marked as “Failed Over”
  3. Set the pChat Active servers in secondary datacenter – at this point we want to make sure all the pChat servers in the secondary datacenter are active.
    1. Set-CsPersistentChatActiveServer –Identity “global” –ActiveServers @{Add=”pchatserver6.domain.com”}2

1 In this scenario if we have active pChat servers split between the datacenters pChat functionality would continue to work. For optimal performance you would still want to follow the steps above in order to failover the backend pChat DB. This should be done during off hours so that production user impact is minimized.

2 Add all pChat servers (using FQDN) in secondary datacenter (West datacenter) separated by commas in between quoted server names.

Once the failover is complete EEPool1 will reconnect the users to the pChat servers located in the secondary (paired) datacenter.

 

Scenario 3B: pChat Pool Server failback

Once the pChat servers in the East datacenter come back online we should failback1. This should be done during off hours so that production user impact is minimized.

  1. Failback to Primary DB with current data – in order to accomplish this follow steps 2-5 from Scenario 2B from the complete site failback above.
  2. Set pChat Pool state as Normal
    1. Set-CsPersistentChatState –Identity “PersistentChatServer:pchatpool1.domain.com” –PoolState Normal
    2. Use the Get-CSPersistentChatState to verify it is marked as “Normal”
  3. Set the pChat Active servers -
    1. Set-CsPersistentChatActiveServer –Identity “global” –ActiveServers @{Add=”pchatserver1.domain.com”}2
  4. pChat clients should now reconnect to the Active pChat servers
  5. Setup Log shipping again to the secondary datacenter for DR purposes
    1. Follow steps here - http://technet.microsoft.com/en-us/library/jj204653.aspx

1 The pChat services will continue to function in this failed over state ever after the pChat servers in the East Datacenter come back online. The reason we want to failback is so that we are in the same production state we were prior to the pChat server failures in the East Datacenter. This will ensure we are in a "Normal" state instead of "Failed Over".

2 Add all pChat servers (using FQDN) that should become active separated by commas in between quoted server names.

 

Scenario 4A: pChat SQL Data Loss

The final failure scenario includes the loss of data from the backend SQL pChat DB (MGC) or User error (deletion). This database includes the pChat room content, principals, and access permissions for the pChat rooms. The pChat data can be backed up in one of the following two ways:

  1. SQL Server DB Backup – backups to SQL databases is usually a standard process for most organizations.
  2. Export-CsPersistentChatData cmdlet – this cmdlet will export all of the pChat data into a .zip file which will contain multiple .xml files (Figure 4)

Figure 4: Export-CsPersistentChatData cmdlet & ZIP file contents

Data that is created by using SQL Server backup requires significantly more disk space—possibly 20 times more—than that created by Export-CsPersistentChatData, but SQL Server backup is more likely to be a procedure that administrators are familiar with. The export in Figure 4 utilizing the Export-CsPersistentChatData cmdlet totaled 150 KB vs. a 4 MB SQL backup.

 

Scenario 4B: pChat SQL Data Recovery

In order to restore the data that we backed up above you can perform the steps below.

  1. If you used SQL Server backup procedures, you must use SQL Server restore procedures.
  2. If you used the Export-CsPersistentChatData cmdlet to back up Persistent Chat data, then you must use the Import-CsPersistentChatData cmdlet to restore the data.

Hopefully this helps everyone understand what process they will need to perform based on their specific failure.

Edge Failover Process - http://technet.microsoft.com/en-us/library/jj721897.aspx

                                             http://technet.microsoft.com/en-us/library/jj688023.aspx

 

 

Read the original blog entry...

More Stories By Richard Schwendiman

My name is Richard Schwendiman and I am currently working for Microsoft as a (PFE) Premier Field Engineer specializing in both Exchange and Lync. I have been working as an IT Consultant for 13+ years focusing on a wide array of Infrastructure technologies. These technologies include Messaging, UC, Networking, Platforms, Active Directory, Virtualization, etc... I am currently certified as an MCSE (Microsoft Certified Systems Engineer), MCSE Messaging 2013, MCSE Communications 2013, MCSA 2012, MCITP Enterprise Messaging, MCTS-Lync, CCNA (Cisco Certified Network Associate), Commvault, CCNP (Cisco Certified Network Professional), and JNCIA-ER (Juniper Enterprise Routing). I am hoping that through this blog I can bring knowledge from the field and keep everyone informed about our ever changing Industry. Please feel free to email me any questions, comments, or concerns pertaining to this blog or any technology related things. Thanks and look forward to providing some good content. http://blogs.technet.com/b/rischwen/

Latest Stories
What sort of WebRTC based applications can we expect to see over the next year and beyond? One way to predict development trends is to see what sorts of applications startups are building. In his session at @ThingsExpo, Arin Sime, founder of WebRTC.ventures, discussed the current and likely future trends in WebRTC application development based on real requests for custom applications from real customers, as well as other public sources of information.
Your homes and cars can be automated and self-serviced. Why can't your storage? From simply asking questions to analyze and troubleshoot your infrastructure, to provisioning storage with snapshots, recovery and replication, your wildest sci-fi dream has come true. In his session at @DevOpsSummit at 20th Cloud Expo, Dan Florea, Director of Product Management at Tintri, provided a ChatOps demo where you can talk to your storage and manage it from anywhere, through Slack and similar services with...
The financial services market is one of the most data-driven industries in the world, yet it’s bogged down by legacy CPU technologies that simply can’t keep up with the task of querying and visualizing billions of records. In his session at 20th Cloud Expo, Karthik Lalithraj, a Principal Solutions Architect at Kinetica, discussed how the advent of advanced in-database analytics on the GPU makes it possible to run sophisticated data science workloads on the same database that is housing the rich...
DevOps at Cloud Expo, taking place October 31 - November 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA, is co-located with 21st Cloud Expo and will feature technical sessions from a rock star conference faculty and the leading industry players in the world. The widespread success of cloud computing is driving the DevOps revolution in enterprise IT. Now as never before, development teams must communicate and collaborate in a dynamic, 24/7/365 environment. There is no time to w...
SYS-CON Events announced today that Massive Networks will exhibit at SYS-CON's 21st International Cloud Expo®, which will take place on Oct 31 – Nov 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA. Massive Networks mission is simple. To help your business operate seamlessly with fast, reliable, and secure internet and network solutions. Improve your customer's experience with outstanding connections to your cloud.
In the enterprise today, connected IoT devices are everywhere – both inside and outside corporate environments. The need to identify, manage, control and secure a quickly growing web of connections and outside devices is making the already challenging task of security even more important, and onerous. In his session at @ThingsExpo, Rich Boyer, CISO and Chief Architect for Security at NTT i3, discussed new ways of thinking and the approaches needed to address the emerging challenges of security i...
"We want to show that our solution is far less expensive with a much better total cost of ownership so we announced several key features. One is called geo-distributed erasure coding, another is support for KVM and we introduced a new capability called Multi-Part," explained Tim Desai, Senior Product Marketing Manager at Hitachi Data Systems, in this SYS-CON.tv interview at 20th Cloud Expo, held June 6-8, 2017, at the Javits Center in New York City, NY.
There is a huge demand for responsive, real-time mobile and web experiences, but current architectural patterns do not easily accommodate applications that respond to events in real time. Common solutions using message queues or HTTP long-polling quickly lead to resiliency, scalability and development velocity challenges. In his session at 21st Cloud Expo, Ryland Degnan, a Senior Software Engineer on the Netflix Edge Platform team, will discuss how by leveraging a reactive stream-based protocol,...
FinTechs use the cloud to operate at the speed and scale of digital financial activity, but are often hindered by the complexity of managing security and compliance in the cloud. In his session at 20th Cloud Expo, Sesh Murthy, co-founder and CTO of Cloud Raxak, showed how proactive and automated cloud security enables FinTechs to leverage the cloud to achieve their business goals. Through business-driven cloud security, FinTechs can speed time-to-market, diminish risk and costs, maintain continu...
DX World EXPO, LLC., a Lighthouse Point, Florida-based startup trade show producer and the creator of "DXWorldEXPO® - Digital Transformation Conference & Expo" has announced its executive management team. The team is headed by Levent Selamoglu, who has been named CEO. "Now is the time for a truly global DX event, to bring together the leading minds from the technology world in a conversation about Digital Transformation," he said in making the announcement.
In his session at 20th Cloud Expo, Mike Johnston, an infrastructure engineer at Supergiant.io, discussed how to use Kubernetes to set up a SaaS infrastructure for your business. Mike Johnston is an infrastructure engineer at Supergiant.io with over 12 years of experience designing, deploying, and maintaining server and workstation infrastructure at all scales. He has experience with brick and mortar data centers as well as cloud providers like Digital Ocean, Amazon Web Services, and Rackspace. H...
Internet of @ThingsExpo, taking place October 31 - November 2, 2017, at the Santa Clara Convention Center in Santa Clara, CA, is co-located with 21st Cloud Expo and will feature technical sessions from a rock star conference faculty and the leading industry players in the world. The Internet of Things (IoT) is the most profound change in personal and enterprise IT since the creation of the Worldwide Web more than 20 years ago. All major researchers estimate there will be tens of billions devic...
"The Striim platform is a full end-to-end streaming integration and analytics platform that is middleware that covers a lot of different use cases," explained Steve Wilkes, Founder and CTO at Striim, in this SYS-CON.tv interview at 20th Cloud Expo, held June 6-8, 2017, at the Javits Center in New York City, NY.
Everything run by electricity will eventually be connected to the Internet. Get ahead of the Internet of Things revolution and join Akvelon expert and IoT industry leader, Sergey Grebnov, in his session at @ThingsExpo, for an educational dive into the world of managing your home, workplace and all the devices they contain with the power of machine-based AI and intelligent Bot services for a completely streamlined experience.
With tough new regulations coming to Europe on data privacy in May 2018, Calligo will explain why in reality the effect is global and transforms how you consider critical data. EU GDPR fundamentally rewrites the rules for cloud, Big Data and IoT. In his session at 21st Cloud Expo, Adam Ryan, Vice President and General Manager EMEA at Calligo, will examine the regulations and provide insight on how it affects technology, challenges the established rules and will usher in new levels of diligence...