Wednesday, 26 June 2013

So why did we call it SALT MINER?


We have announced today with AZ, Roche and Genentech the launch of a consortium to process data from all three companies using Matched Molecular Pair Analysis to extract the medicinal chemistry knowledge. The Grand Rule Database will be at the heart of MedChemica's business and we aim to help many people discovery new drugs for the clinic. We call it SALT MINER - now why did we call it that?
We do live in a world of brands and logos. If you listen to people talk they do like to label things to make them easier to talk about and refer to – just think of the nick names we give each other. So we felt compelled to call our Knowledge sharing consortium using Matched Molecular Pair Analysis something but KSCuMMP doesn't really work. We also work in a world of chem.-informatics and computation and there has been a long history of finding an amusing acronym for the new algorithm you have just written. Just have a good look at the latest edition of Journal of Chemical Information and Modelling.
I was sat in a café (don’t these stories all start like this), with a piece of paper having written out all the words for our complex process of statistical analysis, process, database and exploitation trying to get something to work. I stared blankly at the salt and pepper pots – well there we go. First versions were SALT and CHIPS, SALT and PEPPER, just SALT for while.  We as computational chemists mine data and process it – suddenly SALT MINE. This started to have a good feel because we are based in Cheshire UK. Here there are huge underground salt mines. As these are deep and dry they're now used for long term document storage – lots of information just sitting there….hmmm. Where we used to work "get it up from the salts mines" was the slang term for ordering a report from the archive.
We made it SALT MINER because the system is about exploitation of the knowledge and making better molecules that lead to drug candidates – a person will use the system: this worked.
We do use a lot of Statistical analysis, Large pharma is contributing datasets (SALT) and we will exploit it through Medchemica’s INtegrated design EngineeR. Real salt mining is a tough old business and although what we do is nothing like that level of manual work, the extraction of the nuggets of new knowledge from large datasets is a distinctly non-trivial exercise. Intellectual SALT Miners we are.

Monday, 20 May 2013

Fragment hopping, mapping waters in SBDD and Statistics


So the lecture went well and it's now online at http://lanyrd.com/2013/eurocup6/schbhf/ with all the others at EuroCUP at http://lanyrd.com/2013/eurocup6/.  Which as ever was a scientific delight.  It wasn't a large meeting, but good discussions and interactions throughout the three days. I really like to go to conferences and and be prompted to look at things in a new way. I particularly liked Andreas Evers (Sanofi) approach to combining synthetic accessibility with 3D fragment hopping (it’s been recently published in J.Med Chem too at  http://dx.doi.org/ 10.1021/jm400404v ). 
I learnt a lot more about water mapping and how it can inform structure based drug design from Matt Geballe(Openeye) and Jasna Klicic (Boehringer-Ingleheim), and finally Pat Waters (Vertex) exhortation to use the right statistics more in evaluating methods is a subject close to my heart so I just enjoyed his talk, and excellent slides.

Wednesday, 15 May 2013

Interesting group(s) that increase stability

I am always interested in a functional group that improves metabolic stability of small molecules in med chem (most of publications have been in the area). And so I was intrigued by a small study from Novartis chemists/PK people where they found the cyclopropyl-trifluoromethyl (cp-CF3 {C1(CC1)C(F)(F)F}) was a good replacement for tert-butyl (t-butyl {C(C)(C)C}) - publication here. In a small series of compounds the group increased stability as measured by in-vitro incubation with Rat liver microsomes (RLM) and human liver microsomes (HLM). In one example the increase in stability in RLM was dramatic - worth a go if you what you really need is a tool compound for in-vivo study. The matched pairs of compounds overall showed a reproducible increase in stability (we can recommend converting to log10 to place the data in a linear scale then the fold changes are easier to see). Results correlated to lower clearance in-vivo i.v. dosing. MedChemica's favourite replacement was in there, the iso-butyl nitric (i-bu-CN {C(C)(C)C#N}) and in one matched pair showed an increase in stability too - this group has the advantage of reducing lipophilicity as the same time which may be useful for other properties. The cp-CF3 group is interesting; a CF3 is about the same volume at an iso-propyl so it will be like cp-iPr - however would it be as lipophilic or would the strong pull of the CF3 and Walsh orbit interactions with cp make for an interesting dipole? Do they have measured matched pairs for logD? The group disclosed the synthetic route which involved reaction with diazomethane (I used to make loads of this in my PhD, it is alright to work with but it would put off many. There is an process chemistry route to one of the HIV drugs that so it is possible on scale). I suggest that the reagent to consider making and experimenting with would be (HO)2B-c-Pr-CF3 [C1(CC1)(C(F)(F)F)B(O)O] - a boronic acid of the cyclopropyl. Cyclopropyl boronic acid itself undergoes Susuki reactions really well so I imagine it would. Can you deprotonation Cp-CF3?

Monday, 13 May 2013

Evidence Based Medicinal Chemistry

Giving lectures is a mixed pleasure for me.  Creating a story that is engaging and gets the right messages across takes significant time, especially as I try to avoid “death by bulletpoint” and like to create original graphics. I’ve just finished a lecture for the Openeye EuroCUP meeting (http://www.eyesopen.com/events/eurocup-2013) later this week, where I’m talking on “Towards Evidence Based Medicinal Chemistry via Big Data Analysis”. It’s the antithesis of the sort of “Rule of X” approach, which although performed a great service a decade and a half ago in reducing the number of shear daft compounds that were made, has probably run it’s course as it just doesn’t improve drug hunting enough.

Our approach is akin to that used by the Sky Cycling team’s “aggregation of marginal gains” philosophy. In our case using Matched Pairs based data mining to find the specific changes to make to a molecule that will improve it. The “simple rule approach” gives you things like: “for reducing affinity to the hERG ion channel if you have a base in the molecule, keep the logP down, otherwise, OK”.  Although this is elegant in it’s simplicity, it’s actually not a lot of help.  Simplicity can be a distraction, indeed As Einstein said, "If you want elegance, find a tailor".  Probably another target for me, but not a scientific one. The other more technical reason for not liking the simple rule approaches is based on a beautiful piece of analysis done by Andrew Leach when he was at AstraZeneca and recently published. (Med. Chem. Commun., 2012,3, 528-540 http://dx.doi.org/10.1039/C2MD20010D) where they compared various biological properties of enantiomers of compounds, because enantiomers have the same logP, numbers (and strengths) of hydrogen bonding groups, numbers of aromatic rings, numbers of sp3 centres – all the “simple rule” achiral properties”, and then looked to see how much the enantiomers varied. The enantiomers varied rather a lot – which they shouldn’t do if the drivers of the biological effects were simple achiral properties.  We’re more interested in finding the specific structural changes to a base that might improve it, for example adding an alcohol beta to a basic amine centre or converting a piperidine to a morpholine.
It takes a lot of data and a lot of mining to expose some of these nuggets of knowledge, it’s an area where we have significant experience and are still developing the methods, but we’ve case studies where it’s been paying off. Now, off to EuroCUP.

Friday, 10 May 2013

BioTechs moving into Alderley Park

I'm very pleased to hear the news that three companies have moved into Alderley Park; Blueberry Thera, RedX and Imagen Biotech. AstraZeneca's plan to convert AP into an incubator is starting to pay off - the new reports say the site is called BioHub and will be managed by BioCity Nottingham - there's a brand there.... This is a great boast to the NorthWest and I really hope this keeps on building. Most importantly that the people in the these companies talk, collaborate and innovate together. I heard rumours that there is quite a long list of other companies lining up to move in as well. Good luck to all.

http://www.genengnews.com/gen-news-highlights/az-giving-way-to-early-stage-biopharmas-in-cheshire/81248329/

Monday, 15 April 2013

Matched Molecular Pair Analysis in Drug Discovery


http://www.sciencedirect.com/science/article/pii/S1359644613000937

I'm very pleased to see our review of Matched Molecular Pair Analysis (MMPA) in Drug Discovery in print (link above). We were invited to write this article just at the start of building our new business around Knowledge Based Design - the application of the output from MMPA. We are grateful to the editor for giving us this opportunity and have tried to fill the gap in literature between the review from Ed and Andrew (Graeme Robb and Dan Warner too) back in 2010, but more importantly highlight where the science needs to go. There is so much to do - the science produces a wealth of information, almost too much to take in by the medicinal chemist (or design team) trying to make a decision - it needs simplifying or conversion to visual means (we have a plan - trust us!). But the biggest area to work on it the context problem - we drive this as chemists by alway thinking about 'adding' a group, increasing size - changing a hydrogen (or unsubstituted position) into a group, and tend to think of fluoro, methyl, methoxy, cyano - what about changing your existing group into something else - for example di-methyl amino into cyclopropyl-methyl amino? Now that will improve your hERG! We illustrate this example in the paper. As a result of these, we think we don't have enough date and more specific changes like the context of environment of the group (e.g. electron poor or rich aromatic ring). 'Big Data' we believe is the answer - the more data we can put in and analyse, by the algorithms to auto find match pairs, the more statistically robust the context dependant chemical design rules we can find. With Dan Warners and Steve St. Galleys WizePairZ this will capture information out to four atoms in all directions SO take a chloro replaced by nitrile on an aromatic ring. With four atoms out we could distinguish between a phenyl ring and a pyridine ring which could make all of the difference (read the paper Papadatos/Gillet describe this really well with group examples). So 'Big Data', we make the case for the pooling of data from big pharma to bring together enough to make this happen....

Thursday, 11 April 2013

Predicting Human Serum Albumin binding


http://pubs.acs.org/doi/pdf/10.1021/ci3006098

I have enjoyed reading this morning the publication of Hall, Jorgensen and Whitehead on Automated Ligand- and Structure-Based Protocol for in Silico Prediction of Human Serum Albumin Binding (J. Chem. Inf. Model, ASAP, DOI:10.1021/ci3006098). I have battled a couple of times with acidic compounds which had high albumin binding. This paper was a good summary of several of the approaches to QSAR models and combinations docking and the tricky problem of site selection (3 principle bind site, 3 others and a couple of structures with another 2 - 8 in total). I am pleased they have made the work flow available at a KNIME module with Schroedinger plug ins. They used Induced Fit Docking (IFD) to place molecules into the respective sites with some success by the look of it. I must admit docking Warfarin in does look a nightmare with two lysines and a arginine to bind to.
As we are interested in Knowledge Based Design from Matched Molecular Pair Analysis (MMPA) these sort of QSAR models and docking approaches have a synergy. Can the output of MMPA, the power to suggest molecules to make that will improve a property, be combined with a prediction to improve further the decision of what to make next with high probability of what the make next?